Compare commits

..
Author SHA1 Message Date
renovate-bot 0c74ad09ae chore(deps): update renovate/renovate docker tag to v44.104.0
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 2s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 2s
ci / validate (pull_request) Successful in 2s
renovate-ci / validate-renovate (pull_request) Successful in 7s
ci / build (push) Has been skipped
ci / build (pull_request) Has been skipped
ci / deploy-userbot-panel (push) Has been skipped
ci / deploy-userbot-panel (pull_request) Has been skipped
2026-09-20 22:18:00 +00:00
433 changed files with 25980 additions and 9156 deletions

No files matched your search

-88
View File
@@ -1,88 +0,0 @@
# Build and deployment workflows
Gitea Actions checks this repository, builds its custom images, and deploys
selected services to the workstation. Workflows use the self-hosted runner labels
`linux`, `arch`, and `homelab`; deployment jobs also require `prod`.
## Checks
`ci.yaml` runs Compose validation, actionlint, ShellCheck, Prettier, Ruff,
yamllint, hadolint, and kubeconform. Tool versions are pinned in
`workflows/tool-versions.env` and installed by `install-ci-tools.sh`.
Compose CI checks structure without resolving local environment files or paths.
On the reviewed main commit it only discovers standard filenames; the
`fix/deploy-validation` branch adds the manual Compose entry points too.
Kubeconform validates known resource schemas. Unknown CRDs are skipped. On main,
CI also attempts server-side dry-runs for marked services; these require an
existing namespace and contact the cluster's admission webhooks. A cluster that
is unreachable produces a warning and skips that CI pass. Deploy validation has
its own dry-run stage.
`renovate-ci.yaml` validates Renovate settings and checks that its generated
ConfigMap matches `renovate/renovate.json`.
## Image builds
CI builds changed custom images for `errorpages`, both `homepages` variants, and
the two `edu_master` Python services. Main builds publish `main`, `prod`, and a
commit tag. Dev builds publish `dev`. Build jobs wait for the lint and manifest
checks.
Kubernetes deployment resolves the lab's own registry images to digests, preferring
commit-specific tags. Third-party image versions remain declared in the manifests.
## Deploy selection
`workflows/deploy-lib.sh` owns the stage logic; `ssh-run.sh` invokes it on the
workstation through SSH. Kubernetes selection uses `k8s/active`; Compose selection
uses an `active` file beside a standard `compose.yaml` or `compose.yml`.
Kustomize overlays are supported, although the current tree primarily contains
plain manifests.
Secret files, examples, Helm values, and patch files are excluded from plain
manifest selection. Create local Kubernetes Secrets separately in their target
namespaces. The Helm table lists Prometheus, Loki, Alloy, and Reloader, with each
release controlled by its configured marker. Other charts need separate setup.
## Trigger and required settings
Automatic deployment follows a successful main CI run when the repository Actions
variable `AUTODEPLOY` is `true`. The manual deploy workflow bypasses that switch
and targets the fetched main branch when no validated commit SHA is provided.
A manual dispatch does not prove that this commit passed CI.
Configure the Actions secrets `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_SSH_KEY`, and,
where needed, `DEPLOY_PORT` and `DEPLOY_PATH`. Registry publishing uses
`REGISTRY_USERNAME` and `REGISTRY_PASSWORD`. The remote user needs access to Git,
Docker, kubectl, Helm, jq, and the state directory used for snapshots.
Keep `APPLY_PRUNE` false on the reviewed implementation: its per-file prune loop
is unsafe. `fix/deploy-prune-guard` rejects that option before changes are applied.
Preflight fetches and resets the remote checkout. It refuses when tracked files
have local changes; ignored local env and Secret files stay in place. Do not use
a development checkout with uncommitted tracked changes as the deployment target.
## Stages and recovery
1. Preflight fetches the target commit and checks the remote working tree.
2. Validate selects services, parses Compose, performs Kubernetes dry-runs, and
checks referenced Secrets.
3. Apply Kubernetes records a workload snapshot, upgrades selected Helm releases,
applies resources, and refreshes owned custom images.
4. Apply Compose recreates marked stacks and checks container state.
5. Verify Kubernetes checks changed workloads and attempts rollback for failures.
6. Smoke probes public routes after verification.
The two apply jobs share a remote lock. Workflow concurrency queues deployments
rather than interrupting an older apply. Snapshots live under
`$XDG_STATE_HOME/homelab-deploy`, or `~/.local/state/homelab-deploy` by default.
They contain the pre-apply workload data and commit identifier.
Rollback uses workload revisions. It does not restore ConfigMaps, Secrets,
database schemas, or data. Helm-owned workloads are handled through the Helm
upgrade's rollback path; the generic rollback skips them. Compose has no automatic
rollback. See the [review](../docs/repository-review.md) for remaining recovery
limitations, including SSH retries and serial rollback timing.
-10
View File
@@ -1,10 +0,0 @@
# actionlint configuration. Passed explicitly from the ci workflow:
# actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
#
# The self-hosted act_runner registers custom labels that actionlint cannot know
# about, so declare them here instead of silencing the whole runner-label check.
self-hosted-runner:
labels:
- arch
- homelab
- prod
+125 -305
View File
@@ -7,12 +7,6 @@ on:
pull_request:
workflow_dispatch:
# Every job here is checkout plus local tools. The token needs to read the tree
# and nothing else, and saying so keeps a future step that reaches for the API
# from quietly holding a token that can write to the repository.
permissions:
contents: read
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
@@ -21,87 +15,8 @@ env:
REGISTRY: gcr.forust.xyz
jobs:
lint-compose:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# Structure check for every committed Compose file, active or not.
# Interpolation, env-file and bind-mount resolution are all switched off,
# because inactive stacks have no .env here and would only fail on their
# ${VAR:?} guards. Active stacks get the full check with interpolation in
# the deploy workflow, where the real .env files live.
- name: Validate Compose files
shell: bash
run: |
set -euo pipefail
source .gitea/workflows/compose-lint.sh
mapfile -t safe_flags < <(compose_safe_flags)
echo "docker compose config ${safe_flags[*]-}"
mapfile -t files < <(compose_files)
if [ "${#files[@]}" -eq 0 ]; then
echo "No Compose files found."
exit 0
fi
failed=0
for f in "${files[@]}"; do
if ! out="$(validate_compose_file "$f" ${safe_flags[@]+"${safe_flags[@]}"} 2>&1)"; then
failed=1
echo "::error file=${f}::$(printf '%s' "$out" | head -1)"
fi
done
if [ "$failed" -ne 0 ]; then
echo "Compose validation failed."
exit 1
fi
echo "checked ${#files[@]} Compose file(s)"
lint-actionlint:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint Gitea Actions workflows with actionlint
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint)"
export PATH="$tools_dir:$PATH"
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
lint-shellcheck:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint shell scripts with ShellCheck
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck)"
export PATH="$tools_dir:$PATH"
mapfile -t scripts < <(
git ls-files '*.sh' ':(glob)**/*.bash'
)
if [ "${#scripts[@]}" -eq 0 ]; then
echo "No shell scripts found."
exit 0
fi
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
lint-prettier:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
@@ -109,10 +24,6 @@ jobs:
- name: Check formatting with Prettier
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
export PATH="$tools_dir:$PATH"
mapfile -t prettier_files < <(
git ls-files \
| grep -E '\.(md|json|ya?ml|html|css)$' \
@@ -128,23 +39,17 @@ jobs:
lint-ruff:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint and format-check Python with Ruff
- name: Lint Python with Ruff
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff)"
export PATH="$tools_dir:$PATH"
ruff check .
ruff format --check .
lint-yaml:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
@@ -152,10 +57,6 @@ jobs:
- name: Lint YAML syntax
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
export PATH="$tools_dir:$PATH"
mapfile -t yaml_files < <(
git ls-files '*.yaml' '*.yml' \
':!node_modules/**' \
@@ -171,7 +72,6 @@ jobs:
lint-dockerfiles:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
@@ -179,10 +79,6 @@ jobs:
- name: Lint Dockerfiles
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
export PATH="$tools_dir:$PATH"
mapfile -t dockerfiles < <(
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
)
@@ -196,18 +92,13 @@ jobs:
validate:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 20
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Validate Kubernetes manifests against JSON schemas
- name: Validate Kubernetes manifests
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
mapfile -t manifests < <(
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
@@ -224,102 +115,10 @@ jobs:
-summary \
"${manifests[@]}"
# kubeconform has no schemas for CRDs, so every IngressRoute, Certificate,
# PrometheusRule, Middleware, ServersTransport and ServiceMonitor is silently
# skipped above. The live API server knows the real CRD schemas (and runs the
# cert-manager / Traefik admission webhooks), so validate there too.
#
# Only services marked with a k8s/active marker are checked: server-side
# dry-run needs the target namespace to exist, and inactive services are not
# deployed. Services being enabled for the first time are still covered by
# the JSON-schema pass above.
#
# Main pushes only. `--dry-run=server` persists nothing, but it does execute
# the admission webhooks of the production API server, so anyone able to open
# a pull request would be able to run arbitrary manifest content through
# cert-manager and Traefik. A pull request has nothing to gain from it either:
# only main is ever deployed, and this job runs to completion before the
# deploy workflow is allowed to start, so a bad CRD is still caught before
# anything reaches the cluster -- just on the push rather than on the PR.
- name: Note the server-side check is not running here
if: github.event_name == 'pull_request' || github.ref != 'refs/heads/main'
shell: bash
run: |
echo "::notice::Skipping the server-side dry-run. It executes the cert-manager and" \
"Traefik admission webhooks against the production API server, so it is limited" \
"to pushes to main. CRDs are still schema-checked by kubeconform above, and the" \
"server-side pass still runs on main before the deploy."
- name: Validate active manifests against the live API server
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
shell: bash
run: |
set -euo pipefail
if ! kubectl get --raw='/readyz' --request-timeout=10s >/dev/null 2>&1; then
echo "::warning::Cluster unreachable — skipped server-side validation of CRDs (IngressRoute, Certificate, PrometheusRule). Review manifest changes manually."
exit 0
fi
mapfile -t k8s_dirs < <(
git ls-files '*.yaml' '*.yml' \
| grep -E '(^|/)k8s/' \
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
| sort -u
)
manifests=()
kustomize_apps=()
for dir in "${k8s_dirs[@]}"; do
if [ ! -f "${dir}/active" ]; then
echo "skip (no k8s/active): ${dir}"
continue
fi
if [ -f "${dir}/overlays/prod/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/overlays/prod")
elif [ -f "${dir}/base/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/base")
else
while IFS= read -r f; do
[ -n "$f" ] && manifests+=("$f")
done < <(
git ls-files "${dir}/*.yaml" "${dir}/*.yml" \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
)
fi
done
echo "server-side dry-run: ${#manifests[@]} manifests, ${#kustomize_apps[@]} kustomize apps"
failed=0
for m in ${manifests[@]+"${manifests[@]}"}; do
if ! out="$(kubectl apply --dry-run=server -f "$m" 2>&1)"; then
failed=1
echo "::error file=${m}::$(printf '%s' "$out" | head -1)"
fi
done
for k in ${kustomize_apps[@]+"${kustomize_apps[@]}"}; do
if ! out="$(kubectl apply -k "$k" --dry-run=server 2>&1)"; then
failed=1
echo "::error file=${k}::$(printf '%s' "$out" | head -1)"
fi
done
if [ "$failed" -ne 0 ]; then
echo "Server-side validation failed. The API server (or an admission webhook) rejected these manifests."
exit 1
fi
echo "server-side dry-run: all active manifests accepted by the API server"
build:
needs:
# The panel's scan-deps/test-backend/test-frontend jobs gated here until
# userbot moved to its own repo; upstream's code is upstream's gate now.
# The rule is unchanged: publishing and passing the checks are the same
# gate, so a commit that fails any of these still cannot move :prod.
[lint-actionlint, lint-shellcheck, lint-compose, lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev') && !startsWith(github.ref_name, 'renovate/')
needs: [lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev')
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 60
outputs:
services: ${{ steps.services.outputs.services }}
steps:
@@ -332,20 +131,12 @@ jobs:
id: services
shell: bash
run: |
set -euo pipefail
base="${{ github.event.before }}"
if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then
base="$(git rev-list --max-parents=0 HEAD)"
fi
# A failed diff used to leave changed_files empty, which reads exactly
# like "nothing to build": the job went green having built nothing and
# the tag never moved. The status is checked, not assumed.
if ! changed="$(git diff --name-only "$base" "${GITHUB_SHA}")"; then
echo "::error::cannot diff ${base}..${GITHUB_SHA}"
exit 1
fi
mapfile -t changed_files <<<"$changed"
mapfile -t changed_files < <(git diff --name-only "$base" "${GITHUB_SHA}")
services=()
@@ -365,9 +156,15 @@ jobs:
for file in "${changed_files[@]}"; do
case "$file" in
dtek_notif/*)
add_service dtek_notif
;;
errorpages/*)
add_service errorpages
;;
userbot/*)
add_service userbot
;;
homepages/*)
add_service homepages
;;
@@ -387,63 +184,55 @@ jobs:
echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT"
- name: Log in to registry
# The pin step below also writes (manifest PUTs), and it runs on every
# main push — including manifest-only ones where services is empty. A
# stale persistent login on the old runner used to mask this; a clean
# runner pushes anonymously and gets 401.
if: steps.services.outputs.services != '' || github.ref_name == 'main'
if: steps.services.outputs.services != ''
shell: bash
# Through env, not by substitution into the script. A secret written
# into a run: block is pasted into the shell source before bash parses
# it, so a password containing a quote, a backtick or $(...) becomes
# code that runs. Masking the value in the log does not prevent that.
env:
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: |
set -euo pipefail
printf '%s' "$REGISTRY_PASSWORD" | docker login "${REGISTRY}" \
-u "$REGISTRY_USERNAME" \
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login "${REGISTRY}" \
-u "${{ secrets.REGISTRY_USERNAME }}" \
--password-stdin
- name: Build and push changed images
if: steps.services.outputs.services != ''
shell: bash
run: |
# This step was the one run: block in the workflow without it, and it
# is the one that cannot afford it: a docker push that failed partway
# through the loop used to be followed by more pushes, the loop's exit
# status came from the last one, and the job went green with half the
# images missing from the registry.
set -euo pipefail
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
# Tags for this push. The commit-pinned name is the point of this
# step: the deploy resolves it in preference to :prod, so a deploy
# that sat in the queue behind a later push still gets the build of
# the commit CI validated, instead of whatever :prod points at by the
# time it runs. See render_pinned in deploy-lib.sh.
commit_tag=""
if [ "${GITHUB_REF_NAME}" = "main" ]; then
commit_tag="sha-${GITHUB_SHA:0:12}"
fi
set_tags() {
tags=()
case "${GITHUB_REF_NAME}" in
main) tags+=("main" "prod") ;;
dev) tags+=("dev") ;;
esac
if [ -n "$commit_tag" ]; then
tags+=("$commit_tag")
fi
}
for service in "${services[@]}"; do
case "$service" in
dtek_notif)
image="${REGISTRY}/forust/dtek-notif"
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" dtek_notif
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
errorpages)
image="${REGISTRY}/forust/error-pages"
set_tags
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
@@ -456,17 +245,27 @@ jobs:
docker push "${image}:${tag}"
done
;;
homepages)
for variant in forust xdfnx; do
case "$variant" in
forust)
image="${REGISTRY}/forust/forust-homepage"
userbot)
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
for target in runtime panel; do
case "$target" in
runtime)
context="userbot"
image="${REGISTRY}/forust/userbot"
;;
xdfnx)
image="${REGISTRY}/forust/xdfnx-homepage"
panel)
context="userbot/panel"
image="${REGISTRY}/forust/userbot-panel"
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
@@ -474,15 +273,47 @@ jobs:
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" -f "homepages/Dockerfile.${variant}" homepages
"${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
homepages)
for service in forust xdfnx; do
case "$service" in
forust)
image="${REGISTRY}/forust/forust-homepage"
;;
xdfnx)
image="${REGISTRY}/forust/xdfnx-homepage"
;;
esac
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" -f "homepages/Dockerfile.${service}" homepages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
edu_master)
for variant in session-keeper webinar-checker; do
case "$variant" in
for service in session-keeper webinar-checker; do
case "$service" in
session-keeper)
context="edu_master/phpsessid-bot"
image="${REGISTRY}/forust/session-keeper"
@@ -492,7 +323,15 @@ jobs:
image="${REGISTRY}/forust/webinar-checker"
;;
esac
set_tags
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
@@ -509,41 +348,22 @@ jobs:
esac
done
# Every image the tree names has to carry the commit-pinned name, not only
# the ones this push rebuilt. A push that touches nothing but manifests
# builds nothing, and its deploy would then find no commit-pinned tag to
# resolve and quietly fall back to the moving :prod - which is the whole
# failure the commit-pinned name exists to remove.
#
# Re-tagging copies the manifest list and transfers no layers, so pinning
# six images that already exist costs six registry writes.
#
# The list is derived from the tree rather than written out here, so an
# image added to a manifest is covered without a second place to update.
- name: Pin the commit name on the images this push did not rebuild
if: github.ref_name == 'main'
deploy-userbot-panel:
needs: build
if: github.ref_name == 'main' && contains(needs.build.outputs.services, 'userbot')
runs-on: [self-hosted, linux, arch, homelab, prod]
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Apply and roll out userbot panel
shell: bash
run: |
set -euo pipefail
commit_tag="sha-${GITHUB_SHA:0:12}"
mapfile -t repos < <(
git grep -hoE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+' -- '*.yaml' '*.yml' \
| sort -u
)
if [ "${#repos[@]}" -eq 0 ]; then
echo "No own images referenced by the tree."
exit 0
fi
echo "pinning ${#repos[@]} image(s) to $commit_tag"
for repo in "${repos[@]}"; do
if docker buildx imagetools inspect "$repo:$commit_tag" >/dev/null 2>&1; then
echo " already built by this push: ${repo##*/}"
continue
fi
if ! docker buildx imagetools inspect "$repo:prod" >/dev/null 2>&1; then
echo " WARNING: ${repo##*/} has no :prod to pin and no build produced it"
continue
fi
docker buildx imagetools create --tag "$repo:$commit_tag" "$repo:prod"
echo " pinned ${repo##*/}"
done
kubectl apply -f userbot/k8s/base/panel.yaml
kubectl get secret userbot-common-secrets -n default -o json \
| jq 'del(.metadata.annotations,.metadata.creationTimestamp,.metadata.resourceVersion,.metadata.uid,.metadata.managedFields) | .metadata.namespace = "userbot"' \
| kubectl apply -f -
# Keep legacy deployments (forust/anna) in sync with manifests; they have no replicas field, so apply leaves scaling to the user manager only.
kubectl apply -f userbot/k8s/base/userbots.yaml
kubectl rollout restart deployment/userbot-panel -n userbot
kubectl rollout status deployment/userbot-panel -n userbot --timeout=180s
-46
View File
@@ -1,46 +0,0 @@
#!/usr/bin/env bash
# Shared helpers for validating Compose files. Sourced both by steps in
# .gitea/workflows/ci.yaml and by deploy-lib.sh on the workstation.
#
# Two levels of checking, matching how the repo is structured:
#
# general every committed Compose file, active or not. Pure structure check:
# no ${VAR} interpolation, no .env lookup, no bind-mount path
# resolution. Disabled stacks deliberately have no .env in the repo
# and no values on the CI runner, so a full `config` run would fail on
# their `${VAR:?}` guards for reasons that have nothing to do with the
# change under review.
#
# full active stacks only, with interpolation and env-file resolution, so
# required variables and referenced files are actually resolved. Needs
# the gitignored .env files, so this only runs in the deploy workflow
# on the workstation.
#
# This file is meant to be sourced, not executed.
# All committed Compose files, including the ones deploy never starts.
compose_files() {
git ls-files \
'*/compose.yaml' '*/compose.yml' 'compose.yaml' 'compose.yml' \
'*/docker-compose.yaml' '*/docker-compose.yml'
}
# Prints the flags that turn `docker compose config` into the general check.
# Probed rather than hardcoded so an older Compose without --no-env-resolution
# still gets the flags it does support.
compose_safe_flags() {
local help flag
help="$(docker compose config --help 2>/dev/null || true)"
for flag in --no-interpolate --no-env-resolution --no-path-resolution; do
if printf '%s' "$help" | grep -q -- "$flag"; then
printf '%s\n' "$flag"
fi
done
}
# validate_compose_file <file> [extra docker compose config flags...]
validate_compose_file() {
local file="$1"
shift
docker compose -f "$file" config --quiet "$@"
}
File diff suppressed because it is too large. Load diff
+150 -171
View File
@@ -1,197 +1,176 @@
name: deploy
on:
# Deploy only what CI already validated. workflow_run is used instead of
# workflow_dispatch so a red lint/validate run can never reach the cluster.
workflow_run:
workflows: [ci]
types: [completed]
workflow_dispatch:
# The deploy jobs read the tree, then reach the cluster over SSH with the
# deploy key. The Actions token itself is not part of that path, so it gets
# read-only contents and no more.
permissions:
contents: read
concurrency:
group: deploy-main
# Queue instead of cancelling. Cancelling a run kills the apply job mid-loop and
# takes the verify job down with it, so a superseded deploy would leave the
# cluster half-applied and unchecked — the exact failure the verify job exists
# to catch. kubectl apply and docker compose up are both idempotent, so letting
# the older run finish and then deploying the newer commit costs little.
cancel-in-progress: false
env:
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
# workflow_run's own GITHUB_SHA points at the branch head, not at the commit the
# finished ci run checked. Pin the exact validated commit instead, so a push
# landing mid-deploy cannot make the workstation deploy something else. Also
# what the verify job checks the snapshot against. Empty for workflow_dispatch,
# which falls back to the current origin/main.
DEPLOY_SHA: ${{ github.event.workflow_run.head_sha }}
jobs:
preflight:
# Autodeploy defaults to OFF: pushes deploy only when the AUTODEPLOY repo
# variable is set to 'true' (Settings -> Actions -> Variables). A manual
# Run workflow always bypasses the switch: dispatching it is the explicit
# intent to deploy.
if: >-
(vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') &&
(github.event_name != 'workflow_run' ||
(github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.head_branch == 'main'))
redeploy:
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Fetch and reset workstation
- name: Redeploy workstation
shell: bash
env:
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
# Set APPLY_PRUNE=true to enable kubectl apply --prune. Requires every
# manifest to carry label app.kubernetes.io/managed-by=homelab-deploy,
# otherwise previously applied resources get deleted on the next run.
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh preflight
validate:
needs: [preflight]
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 20
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
- name: Dry-run manifests and check Secrets
shell: bash
run: |
deploy_port="${DEPLOY_PORT:-22}"
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
ssh_key="$RUNNER_TEMP/deploy_key"
mkdir -p "$RUNNER_TEMP"
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
chmod 600 "$ssh_key"
ssh_opts=(
-i "$ssh_key"
-p "$deploy_port"
-o BatchMode=yes
-o StrictHostKeyChecking=accept-new
)
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
"DEPLOY_PATH=$(printf '%q' \"$deploy_path\") APPLY_PRUNE=$(printf '%q' \"${APPLY_PRUNE:-false}\") bash -se" <<'EOF'
set -euo pipefail
./.gitea/workflows/ssh-run.sh validate
apply-k8s:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
# Apply only, no verification, so this is just the work itself: snapshot,
# then sequential `helm upgrade --install --wait --rollback-on-failure --timeout 10m`, then the apply loop.
# Verification has its own job and its own budget.
#
# 45 is roughly four times the measured cost of the stage, which is
# deliberately not raised on a theory:
#
# helm, healthy 3 no-op upgrades ~3-5 min
# helm, one release bad rollback-on-failure spends its 10m, ~10-15 min
# then rolls that one back
# apply loop ~40 manifests, 4 of which ~1 min
# resolve an image digest
# restart_stale_images 7.6s to find 8 workloads, ~0.5 min
# 9.8s to resolve their digests
#
# The helm figure is one release, not three: `set -e` aborts
# upgrade_helm_releases on the first failure, so a broken release costs
# 10m and the other two are never attempted. Multiplying 10m by three
# overstates the worst case by 20 minutes.
#
# The 45 minutes this was last raised to 45 were still not enough, and the
# job logs for those runs no longer exist, so what actually consumed the
# budget is not known - the two measurable candidates above account for
# ~15 of it. The unbounded `docker manifest inspect` against the registry's
# known hang mode is now bounded inside registry_digest (25s timeout, 3
# attempts): a dead registry fails each owned image after ~85s instead of
# hanging the stage, and a blinking one is retried instead of failing the
# whole apply file. Still open: make the stage announce which manifest it
# is working on, so a killed run leaves a diagnosable last line.
timeout-minutes: 45
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
repo="${DEPLOY_PATH:-/srv/homelab}"
- name: Apply Kubernetes manifests
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-k8s
if [ ! -d "$repo/.git" ]; then
echo "Repository not found at $repo"
exit 1
fi
apply-compose:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
git -C "$repo" fetch origin main
git -C "$repo" reset --hard origin/main
- name: Redeploy docker compose stacks
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-compose
# Runtime selection: a service is k8s-managed when $SERVICE/k8s/active
# exists. Otherwise it is compose-managed, and only k8s/routing/*
# manifests (external Services / EndpointSlices / ServersTransport /
# Ingresses that route to docker backends) are applied.
# migrate: touch SERVICE/k8s/active (+ move routing files up)
# rollback: rm SERVICE/k8s/active
collect_k8s() {
find "$1" -type f \( -name '*.yaml' -o -name '*.yml' \) \
! -path '*/routing/*' ! -path '*/overlays/*' \
! -name 'kustomization.y*ml' ! -name '*.example.y*ml' \
! -name '*values.y*ml' ! -name 'patch-*.y*ml' \
| sort
}
# Watches the workloads this deploy changed and rolls back the ones that never
# became healthy. Runs even when the apply jobs failed, timed out or were
# cancelled — that is the whole point of splitting it out. `always()` is what
# lets it start after a failed dependency; the needs on apply-compose are a
# barrier, so verification begins only once both applies are done.
verify-k8s:
needs: [apply-k8s, apply-compose]
if: >-
always() &&
needs.apply-k8s.result != 'skipped' &&
needs.apply-compose.result != 'skipped'
runs-on: [self-hosted, linux, arch, homelab, prod]
# Not raised, because the arithmetic does not close.
#
# 32 workloads are under management and the wave width is 8, so the verify
# itself is 4 waves of ROLLOUT_TIMEOUT (300s) = 20 minutes worst case, when
# every rollout times out rather than converging. That is already 20 of 30.
#
# The other 10 would have to absorb rollback, and rollback_workloads is a
# serial `while read` loop at 300s per failed workload. 10 minutes buys two.
# Any larger number is buying a bigger multiple of an unbounded term rather
# than covering a known cost: 60 minutes buys eight, and 60 minutes is
# therefore not a bound, it is a guess with two digits.
#
# The number becomes derivable the moment rollback uses the same wave width
# as the verify: 32 failures then cost 4 waves = 20 minutes instead of 160,
# and 45 covers verify plus rollback at full width. That change is to the
# recovery path and is not folded into a timeout edit.
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
collect_k8s_inactive() {
find "$1" -type f \( -name '*.yaml' -o -name '*.yml' \) \
\( -name 'namespace.y*ml' -o -path '*/routing/*' \) \
! -path '*/overlays/*' ! -name '*.example.y*ml' \
| sort
}
- name: Verify workloads and roll back on failure
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh verify-k8s
mapfile -t compose_stacks < <(
find "$repo" -type f \( -name 'compose.yaml' -o -name 'compose.yml' \) | sort
)
# Asks the public route of every active service whether it is actually
# serving, which the rollout check above structurally cannot: a pod can
# converge and still be crash-looping, or be listening on a port no Service
# points at, or answer 500.
#
# `always()` for the same reason verify-k8s has it, and it runs after that job
# specifically because a rollback is when a route most needs re-checking. The
# needs is a barrier, not a filter: whether verify-k8s passed, failed or was
# cancelled, the probes are what say whether the cluster is serving, and
# suppressing them on a rollback would hide the one run where the answer
# matters most.
smoke:
needs: [verify-k8s]
if: always() && needs.verify-k8s.result != 'skipped'
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
mapfile -t k8s_manifests < <(
for kd in $(find "$repo" -type d -name k8s ! -path '*/.git/*' | sort); do
if [ -f "$kd/active" ]; then
collect_k8s "$kd"
else
collect_k8s_inactive "$kd"
fi
done
)
- name: Probe the public route of every active service
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh smoke
echo "== Validate compose stacks =="
for cf in "${compose_stacks[@]}"; do
dir=$(dirname "$cf")
if [ -f "$dir/k8s/active" ]; then
echo " skip (k8s-managed): $dir"
continue
fi
echo " config: $cf"
docker compose -f "$cf" config --quiet
done
echo "== Validate k8s manifests (kubectl dry-run) =="
for m in "${k8s_manifests[@]}"; do
echo " apply --dry-run=client $m"
kubectl apply --dry-run=client -f "$m" >/dev/null
done
echo "== Applying Kubernetes manifests =="
ns_files=()
other_files=()
for m in "${k8s_manifests[@]}"; do
case "$m" in
*/namespace.y?ml) ns_files+=("$m") ;;
*) other_files+=("$m") ;;
esac
done
prune_opts=()
if [ "${APPLY_PRUNE:-false}" = "true" ]; then
prune_opts=(--prune -l app.kubernetes.io/managed-by=homelab-deploy)
fi
if [ "${#ns_files[@]}" -gt 0 ]; then
echo " namespaces first: ${ns_files[*]}"
kubectl apply -f "${ns_files[@]}"
fi
if [ -f "$repo/prometheus-stack/k8s/active" ]; then
echo "== Upgrading kube-prometheus-stack =="
helm upgrade --install prometheus-stack prometheus-community/kube-prometheus-stack \
--namespace prometheus \
--version 86.2.3 \
--values "$repo/prometheus-stack/k8s/grafana-values.yaml" \
--wait
fi
if [ -f "$repo/loki/k8s/active" ]; then
echo "== Upgrading loki/alloy =="
helm repo add grafana https://grafana.github.io/helm-charts >/dev/null 2>&1 || true
helm repo update grafana >/dev/null 2>&1 || true
helm upgrade --install loki grafana/loki \
--version 7.3.0 \
--namespace prometheus \
--values "$repo/loki/k8s/loki-values.yaml" \
--wait
helm upgrade --install alloy grafana/alloy \
--version 1.12.1 \
--namespace prometheus \
--values "$repo/loki/k8s/alloy-values.yaml" \
--wait
fi
if [ "${#other_files[@]}" -gt 0 ]; then
echo " resources: ${other_files[*]}"
kubectl apply "${prune_opts[@]}" -f "${other_files[@]}"
fi
echo "== Redeploying docker compose stacks =="
for cf in "${compose_stacks[@]}"; do
dir=$(dirname "$cf")
if [ -f "$dir/k8s/active" ]; then
echo " skip (k8s-managed): $dir"
continue
fi
echo " compose: $dir"
if grep -Eq '^\s+pull_policy:\s*build\b' "$cf"; then
docker compose -f "$cf" build
docker compose -f "$cf" push
fi
docker compose -f "$cf" up -d --pull always --remove-orphans
done
EOF
-254
View File
@@ -1,254 +0,0 @@
#!/usr/bin/env bash
# Installs the pinned CI tools into "$TOOLS_DIR/bin" and echoes that directory
# on stdout, so callers can do:
#
# export PATH="$(bash .gitea/workflows/install-ci-tools.sh kubeconform shellcheck):$PATH"
#
# Versions come from tool-versions.env next to this script and are kept fresh by
# Renovate. Re-running is cheap: an already-installed tool at the pinned version
# is left alone.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=tool-versions.env
. "$here/tool-versions.env"
TOOLS_DIR="${TOOLS_DIR:-${RUNNER_TEMP:-/tmp}/homelab-tools}"
BIN_DIR="$TOOLS_DIR/bin"
mkdir -p "$BIN_DIR"
# The just-installed tools must resolve inside this script too: callers only
# prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would
# miss the binary install_uv just placed (exit 127 on a clean runner).
export PATH="$BIN_DIR:$PATH"
arch="$(uname -m)"
# Upstream projects disagree on arch spelling: kubeconform and actionlint use
# Go names (amd64/arm64), shellcheck uses uname names (x86_64/aarch64), node
# uses neither (x64/arm64), and hadolint mixes the two in a single release
# (x86_64 but arm64).
case "$arch" in
x86_64 | amd64)
goarch=amd64
sharch=x86_64
nodearch=x64
hadolintarch=x86_64
;;
aarch64 | arm64)
goarch=arm64
sharch=aarch64
nodearch=arm64
hadolintarch=arm64
;;
*)
echo "install-ci-tools: unsupported architecture: $arch" >&2
exit 1
;;
esac
fetch() {
# fetch <url> <dest>
if command -v curl >/dev/null 2>&1; then
curl -sSLf --retry 3 -o "$2" "$1"
elif command -v wget >/dev/null 2>&1; then
wget -q -O "$2" "$1"
else
echo "install-ci-tools: neither curl nor wget is available" >&2
exit 1
fi
}
# resolve <command>
# Absolute path to use for invoking a tool: the copy in BIN_DIR when present,
# otherwise the name for PATH lookup. Every version check and every in-script
# invocation goes through this, so a tool missing from both places reads as
# "not installed" instead of dying with 127 under `set -e`.
resolve() {
if [ -x "$BIN_DIR/$1" ]; then
printf '%s' "$BIN_DIR/$1"
else
printf '%s' "$1"
fi
}
# installed_version <command>
# Prints the version of an already-installed tool, or nothing. Each tool spells
# its version flag differently, hence the case.
installed_version() {
local bin out
bin="$(resolve "$1")"
if ! command -v "$bin" >/dev/null 2>&1; then
return 0
fi
case "$1" in
kubeconform) out="$("$bin" -v 2>/dev/null | head -1 || true)" ;;
*) out="$("$bin" --version 2>/dev/null | head -1 || true)" ;;
esac
printf '%s' "$out"
}
# at_version <command> <expected>
at_version() {
case "$(installed_version "$1")" in
*"$2"*) return 0 ;;
*) return 1 ;;
esac
}
install_kubeconform() {
if at_version kubeconform "v${KUBECONFORM_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/yannh/kubeconform/releases/download/v${KUBECONFORM_VERSION}/kubeconform-linux-${goarch}.tar.gz" \
"$tmp/kubeconform.tar.gz"
tar -xzf "$tmp/kubeconform.tar.gz" -C "$tmp" kubeconform
install -m 0755 "$tmp/kubeconform" "$BIN_DIR/kubeconform"
rm -rf "$tmp"
}
install_shellcheck() {
if at_version shellcheck "${SHELLCHECK_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/koalaman/shellcheck/releases/download/v${SHELLCHECK_VERSION}/shellcheck-v${SHELLCHECK_VERSION}.linux.${sharch}.tar.xz" \
"$tmp/shellcheck.tar.xz"
tar -xJf "$tmp/shellcheck.tar.xz" -C "$tmp" --strip-components=1 "shellcheck-v${SHELLCHECK_VERSION}/shellcheck"
install -m 0755 "$tmp/shellcheck" "$BIN_DIR/shellcheck"
rm -rf "$tmp"
}
install_uv() {
if at_version uv "${UV_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
# uv release tags carry no leading v, unlike every other tool installed here.
fetch "https://github.com/astral-sh/uv/releases/download/${UV_VERSION}/uv-${sharch}-unknown-linux-gnu.tar.gz" \
"$tmp/uv.tar.gz"
tar -xzf "$tmp/uv.tar.gz" -C "$tmp" --strip-components=1 "uv-${sharch}-unknown-linux-gnu/uv"
install -m 0755 "$tmp/uv" "$BIN_DIR/uv"
rm -rf "$tmp"
}
install_hadolint() {
if at_version hadolint "${HADOLINT_VERSION}"; then
return 0
fi
# A bare binary, no archive: hadolint ships one file per platform.
fetch "https://github.com/hadolint/hadolint/releases/download/v${HADOLINT_VERSION}/hadolint-linux-${hadolintarch}" \
"$BIN_DIR/hadolint"
chmod 0755 "$BIN_DIR/hadolint"
}
# ruff and yamllint both come from PyPI as wheels, which uv unpacks for us.
install_uv_tool() {
# <package> <pinned version>
if at_version "$1" "$2"; then
return 0
fi
install_uv
UV_TOOL_BIN_DIR="$BIN_DIR" "$BIN_DIR/uv" tool install --force "$1==$2" >/dev/null
}
install_ruff() {
install_uv_tool ruff "${RUFF_VERSION}"
}
install_yamllint() {
install_uv_tool yamllint "${YAMLLINT_VERSION}"
}
install_pip_audit() {
install_uv_tool pip-audit "${PIP_AUDIT_VERSION}"
}
install_prettier() {
if at_version prettier "${PRETTIER_VERSION}"; then
return 0
fi
# Not a standalone binary: prettier's entry point requires ../package.json
# relative to its own real path, so the package directory has to survive
# next to it. Hence a versioned directory plus a relative symlink, rather
# than copying the one file out as the other installers do.
local dir="$BIN_DIR/prettier-${PRETTIER_VERSION}"
if [ ! -f "$dir/package/package.json" ]; then
rm -rf "$dir"
mkdir -p "$dir"
fetch "https://registry.npmjs.org/prettier/-/prettier-${PRETTIER_VERSION}.tgz" "$dir/prettier.tgz"
tar -xzf "$dir/prettier.tgz" -C "$dir"
rm -f "$dir/prettier.tgz"
# npm strips the exec bit from bin/ on the way into the tarball.
chmod 0755 "$dir/package/bin/prettier.cjs"
fi
# Relative, so the whole tree stays valid if TOOLS_DIR is relocated.
ln -sfn "prettier-${PRETTIER_VERSION}/package/bin/prettier.cjs" "$BIN_DIR/prettier"
}
install_node() {
# npm gets checked by running it, not by looking it up: what matters is that
# it answers, so a stub, a half-removed Arch package or a name that resolves
# to something broken all have to read as "not installed". The runner's npm
# is a symlink into /usr/lib/node_modules/npm, which is exactly the kind of
# thing that disappears between runs.
if at_version node "v${NODE_VERSION}" && [ -n "$(installed_version npm)" ]; then
return 0
fi
# Same shape as prettier above: the tarball's bin/npm and bin/npx are links
# into lib/node_modules, so the whole tree has to survive next to them.
local dir="$BIN_DIR/node-${NODE_VERSION}"
if [ ! -x "$dir/bin/node" ]; then
rm -rf "$dir"
mkdir -p "$dir"
fetch "https://nodejs.org/dist/v${NODE_VERSION}/node-v${NODE_VERSION}-linux-${nodearch}.tar.xz" \
"$dir/node.tar.xz"
tar -xJf "$dir/node.tar.xz" -C "$dir" --strip-components=1 "node-v${NODE_VERSION}-linux-${nodearch}"
rm -f "$dir/node.tar.xz"
fi
# Relative, so the whole tree stays valid if TOOLS_DIR is relocated.
for bin in node npm npx; do
ln -sfn "node-${NODE_VERSION}/bin/${bin}" "$BIN_DIR/${bin}"
done
}
install_actionlint() {
if at_version actionlint "${ACTIONLINT_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_${goarch}.tar.gz" \
"$tmp/actionlint.tar.gz"
tar -xzf "$tmp/actionlint.tar.gz" -C "$tmp" actionlint
install -m 0755 "$tmp/actionlint" "$BIN_DIR/actionlint"
rm -rf "$tmp"
}
wanted=("$@")
if [ "${#wanted[@]}" -eq 0 ]; then
wanted=(kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
fi
for tool in "${wanted[@]}"; do
case "$tool" in
kubeconform) install_kubeconform ;;
shellcheck) install_shellcheck ;;
actionlint) install_actionlint ;;
prettier) install_prettier ;;
ruff) install_ruff ;;
yamllint) install_yamllint ;;
pip-audit) install_pip_audit ;;
hadolint) install_hadolint ;;
node) install_node ;;
uv) install_uv ;;
*)
echo "install-ci-tools: unknown tool: $tool" >&2
exit 1
;;
esac
done
printf '%s\n' "$BIN_DIR"
+18 -42
View File
@@ -7,58 +7,33 @@ on:
- main
workflow_dispatch:
permissions:
contents: read
jobs:
validate-renovate:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 20
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag,
# so the same version that runs in the cluster is the one validated here.
- name: Resolve the deployed Renovate image
id: image
- name: Validate Renovate Compose draft
shell: bash
run: |
set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)"
if [ -z "$image" ]; then
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
exit 1
fi
echo "using $image"
echo "image=$image" >> "$GITHUB_OUTPUT"
trap 'rm -f renovate/.env' EXIT
printf '%s\n' \
'RENOVATE_ENDPOINT=https://gitea.example/api/v1' \
'RENOVATE_TOKEN=test-token' \
'RENOVATE_REPOSITORIES=forust/homelab' \
> renovate/.env
docker compose -f renovate/renovate-compose.yaml config --quiet
- name: Validate Renovate repository config
- name: Validate Kubernetes manifests
shell: bash
run: |
set -euo pipefail
docker run --rm \
-v "$PWD/renovate:/opt/renovate:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"${{ steps.image.outputs.image }}" \
renovate-config-validator /opt/renovate/renovate.json
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
# carries an inlined copy of the config. Fail if it no longer matches.
- name: Check the generated Renovate ConfigMap
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/sync-renovate-configmap.sh --check
- name: Validate Renovate Kubernetes manifests
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
kubeconform \
-v "$PWD:/work" \
-w /work \
ghcr.io/yannh/kubeconform:latest \
-strict \
-ignore-missing-schemas \
-summary \
@@ -66,11 +41,12 @@ jobs:
renovate/k8s/configmap.yaml \
renovate/k8s/cronjob.yaml
- name: Validate Renovate Compose file
- name: Validate Renovate repository config
shell: bash
run: |
set -euo pipefail
source .gitea/workflows/compose-lint.sh
mapfile -t safe_flags < <(compose_safe_flags)
validate_compose_file renovate/renovate-compose.yaml \
${safe_flags[@]+"${safe_flags[@]}"}
docker run --rm \
-v "$PWD:/work" \
-w /work \
renovate/renovate:44.103.0 \
renovate-config-validator renovate.json
+7 -30
View File
@@ -21,11 +21,6 @@ on:
default: false
type: boolean
# Renovate writes through its own bot PAT, passed in as RENOVATE_TOKEN, so the
# Actions token is only ever used to read the checkout.
permissions:
contents: read
concurrency:
group: renovate-run
cancel-in-progress: false
@@ -33,36 +28,18 @@ concurrency:
jobs:
run-renovate:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 60
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
# Reading it here means this workflow validates and runs the exact version
# that is deployed, instead of a copy that silently goes stale.
- name: Resolve the deployed Renovate image
id: image
shell: bash
run: |
set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)"
if [ -z "$image" ]; then
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
exit 1
fi
echo "using $image"
echo "image=$image" >> "$GITHUB_OUTPUT"
- name: Validate Renovate config
shell: bash
run: |
set -euo pipefail
docker run --rm \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"${{ steps.image.outputs.image }}" \
-v "$PWD/renovate/config.js:/opt/renovate/config.js:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/config.js \
renovate/renovate:44.103.0 \
renovate-config-validator
- name: Run Renovate
@@ -79,14 +56,14 @@ jobs:
: "${RENOVATE_TOKEN:?missing RENOVATE_TOKEN secret — add a renovate-bot PAT in repo/org Actions secrets}"
docker run --rm \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
-v "$PWD/renovate/config.js:/opt/renovate/config.js:ro" \
-e RENOVATE_PLATFORM=gitea \
-e RENOVATE_ENDPOINT=https://git.forust.xyz/api/v1 \
-e RENOVATE_ENDPOINT=https://gitea.forust.xyz/api/v1 \
-e RENOVATE_TOKEN="$RENOVATE_TOKEN" \
-e RENOVATE_GITHUB_COM_TOKEN="${RENOVATE_GITHUB_COM_TOKEN:-}" \
-e RENOVATE_REPOSITORIES="${RENOVATE_REPOSITORIES:-forust/homelab}" \
-e RENOVATE_DRY_RUN="${RENOVATE_DRY_RUN:-}" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
-e RENOVATE_CONFIG_FILE=/opt/renovate/config.js \
-e RENOVATE_BASE_DIR=/tmp/renovate \
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
"${{ steps.image.outputs.image }}"
renovate/renovate:44.103.0
-71
View File
@@ -1,71 +0,0 @@
#!/usr/bin/env bash
# usage: ssh-run.sh <stage>
# Runs one deploy-lib.sh stage on the workstation over SSH.
set -euo pipefail
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
deploy_port="${DEPLOY_PORT:-22}"
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)"
# The private key is written to a per-run directory that is removed on exit, so a
# failed or cancelled job cannot leave deploy credentials in the runner's temp
# directory. Do not use a fixed path: apply-k8s and apply-compose run in parallel.
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
trap 'rm -rf "$key_dir"' EXIT INT TERM
ssh_key="$key_dir/deploy_key"
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
chmod 600 "$ssh_key"
# A connection that died silently used to hang until the job timeout, and the
# stage was never re-run: one flaky TCP session cost a whole 45-minute apply.
# ServerAlive* bounds how long a dead peer goes unnoticed, ConnectTimeout bounds
# setup. Only exit 255 - ssh's own transport failures - is retried. A stage that
# fails on its own merits exits with the remote's status, so a real failure
# still surfaces its own log instead of burning three attempts. The stages are
# declarative applies, so re-running one that had already committed is harmless.
ssh_opts=(
-i "$ssh_key" -p "$deploy_port"
-o BatchMode=yes -o StrictHostKeyChecking=accept-new
-o ConnectTimeout=15
-o ServerAliveInterval=15 -o ServerAliveCountMax=4
)
rc=0
# apply-k8s and apply-compose are separate workflow jobs so the graph stays
# intact for the verify job, but on a single node they must not run at once:
# host docker churn on top of cluster churn is what melts the node (load 40+,
# netbird/ssh die, helm is left pending-*). Serialize them on the workstation
# with a shared lock; whoever arrives second waits.
remote_cmd=(bash -se)
case "$1" in
apply-k8s | apply-compose)
remote_cmd=(flock -w 5400 /tmp/homelab-apply.lock bash -se)
;;
esac
for attempt in 1 2 3; do
if [ "$attempt" -gt 1 ]; then
echo ":: warning::ssh transport failed, retrying (${attempt}/3)"
sleep $((attempt * 5))
fi
rc=0
# shellcheck disable=SC2029 # remote_cmd/ssh_opts expand on the client on purpose: they select the local ssh invocation, only the heredoc runs remotely.
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
env "REPO=$deploy_path" "APPLY_PRUNE=${APPLY_PRUNE:-false}" \
"DEPLOY_SHA=${DEPLOY_SHA:-}" "DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-}" \
"STAGE=$1" "${remote_cmd[@]}" <<'EOF' || rc=$?
source "$REPO/.gitea/workflows/deploy-lib.sh"
run_stage "$STAGE"
EOF
[ "$rc" -eq 0 ] && break
[ "$rc" -ne 255 ] && break
done
if [ "$rc" -ne 0 ]; then
echo ":: error::stage $1 failed over ssh (exit $rc)"
fi
exit "$rc"
@@ -1,55 +0,0 @@
#!/usr/bin/env bash
# Regenerates renovate/k8s/configmap.yaml from renovate/renovate.json.
#
# renovate/renovate.json is the single source of truth: the CronJob, the Compose
# file and the renovate-run workflow all mount that exact file. A ConfigMap cannot
# read a file from the repository, so the same bytes are inlined here as a literal
# block. This script keeps the copy honest:
#
# .gitea/workflows/sync-renovate-configmap.sh # rewrite in place
# .gitea/workflows/sync-renovate-configmap.sh --check # fail if out of date
#
# renovate-ci runs the --check form on every PR and push, so a config change that
# forgets to regenerate the ConfigMap cannot be merged.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
repo="$(git -C "$here" rev-parse --show-toplevel)"
src="$repo/renovate/renovate.json"
dst="$repo/renovate/k8s/configmap.yaml"
[ -f "$src" ] || {
echo "missing $src" >&2
exit 1
}
render() {
cat <<'HEADER'
# GENERATED FILE - do not edit by hand.
# Source: renovate/renovate.json
# Regenerate: .gitea/workflows/sync-renovate-configmap.sh
# Verify: .gitea/workflows/sync-renovate-configmap.sh --check
apiVersion: v1
kind: ConfigMap
metadata:
name: renovate-config
namespace: renovate
data:
renovate.json: |
HEADER
sed 's/^/ /' "$src"
}
if [ "${1:-}" = "--check" ]; then
if ! diff -u "$dst" <(render) >/dev/null 2>&1; then
echo "ERROR: $dst is out of sync with renovate/renovate.json"
echo "Run: .gitea/workflows/sync-renovate-configmap.sh"
diff -u "$dst" <(render) || true
exit 1
fi
echo "renovate/k8s/configmap.yaml is in sync with renovate/renovate.json"
exit 0
fi
render >"$dst"
echo "wrote $dst"
-33
View File
@@ -1,33 +0,0 @@
# Pinned versions of the CI tools installed by install-ci-tools.sh.
# Renovate keeps these up to date (see customManagers in renovate/renovate.json).
#
# Every version here except NODE_VERSION matches what was already installed on
# the runner, so pinning them changes what CI does not at all. It changes what
# CI does when the runner is rebuilt with something else: today
# install-ci-tools.sh finds the pinned version already on PATH and installs
# nothing, and a runner that drifts gets the pinned one installed over it.
#
# The renovate image version is NOT pinned here: renovate/k8s/cronjob.yaml is the
# single source of truth and the workflows read the tag from it, so there is
# nothing to drift.
ACTIONLINT_VERSION="1.7.7"
SHELLCHECK_VERSION="0.11.0"
KUBECONFORM_VERSION="0.8.0"
PRETTIER_VERSION="3.8.1"
RUFF_VERSION="0.16.8"
YAMLLINT_VERSION="1.38.0"
HADOLINT_VERSION="2.14.0"
# pip-audit reads the advisory database over the network, so a floating version
# would make the same commit report different things on different days. Pin it
# like the rest: the advisories themselves are the moving part, not the tool.
PIP_AUDIT_VERSION="2.10.1"
# uv builds the throwaway venv the pytest job runs in, and unpacks the PyPI
# wheels for ruff, yamllint and pip-audit.
UV_VERSION="0.12.17"
# node runs `npm ci` for the frontend tests and the npm audit, and it is the one
# pin here that does NOT come from the runner: the runner's system node is a
# rolling Arch package (it was node 26 with no npm at all when this was pinned),
# and the panel image is node:22-alpine. Pinned to the image's major on purpose,
# so the tree that gets tested is the tree that gets built. Renovate keeps this
# in step with the Dockerfile's node: tag via the "node runtime" group.
NODE_VERSION="22.23.3"
-5
View File
@@ -21,9 +21,6 @@ checkmk/checkmk/*
downtify/Downtify_downloads
headscale/config/*
headscale/data/*
# NetBird local hostnames and generated secrets
netbird/.env
netbird/secrets/
searxng/core-config/*
# Steaming services files
@@ -96,8 +93,6 @@ replacements.txt
# Temp files
edu_master/temp/
temp/*
# Local-only tooling scratch space (pinned CI tools, verification scripts)
tmp/
# Environment
.env
-163
View File
@@ -1,163 +0,0 @@
# Homelab
Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the
Gitea Actions that build and deploy them. Most applications have both deployment
formats. Headscale, Nextcloud AIO, and the media stack run on Docker; Kubernetes
provides their ingress through Services and EndpointSlices.
These files contain this lab's domains, IP addresses, storage paths, and private
registry names. Running them on another machine takes some editing.
## Start here
- [Service list](#services) — what each directory contains.
- [Deployment workflow](.gitea/README.md) — selection, validation, and recovery.
- [Repository review](docs/repository-review.md) — confirmed problems and fix branches.
- [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and
[cert-manager](cert-manager/README.md) — common dependencies.
## What gets deployed
The `active` files are switches for the deploy workflow, not health indicators.
| File | Effect |
| ---------------------- | ----------------------------------------------------------- |
| `<service>/active` | Include that directory's `compose.yaml` or `compose.yml`. |
| `<service>/k8s/active` | Include its Kubernetes manifests or Kustomize overlay. |
| Both | Run the Compose stack and apply the Kubernetes resources. |
| Neither | Keep the configuration in Git without automatic deployment. |
`shared-compose.yaml`, `client.compose.yaml`, and `renovate-compose.yaml` are
manual entry points. The deploy script does not discover them.
Kubernetes selection excludes secret files, examples, Helm values, and patches.
Helm releases listed in `deploy-lib.sh` are upgraded separately. Traefik,
cert-manager, and CrowdSec have additional bootstrap steps; an `active` marker
does not install their charts.
The table below describes committed configuration. It does not claim that a
service is currently healthy or running.
## Services
| Service | Configuration | Selected by markers |
| ----------------------------------------------------------- | ---------------------------- | ------------------- |
| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes |
| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes |
| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual |
| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes |
| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual |
| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual |
| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes |
| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual |
| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual |
| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual |
| [EDU session keeper and Telegram bot](edu_master/README.md) | Kubernetes + Compose | Kubernetes |
| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes |
| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes |
| [Glance](glance/README.md) | Kubernetes + Compose | Manual |
| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual |
| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes |
| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes |
| [Kener](kener/README.md) | Kubernetes + Compose | Manual |
| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes |
| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes |
| [n8n](n8n/README.md) | Kubernetes + Compose | Manual |
| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes |
| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes |
| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes |
| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Penpot](penpot/README.md) | Compose | Manual |
| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual |
| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes |
| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes |
| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes |
| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes |
| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes |
| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual |
| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Termix](termix/README.md) | Kubernetes + Compose | Manual |
| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes |
| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes |
| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes |
| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes |
## Running a Compose stack
Use the service README first. Where a service has an env example, copy it inside
that service's directory and replace the placeholders. The root `.env.example`
is an older collection of variables, not a complete configuration for every stack.
For example, from the repository root:
```sh
cd netbox
cp .env.example .env
$EDITOR .env
docker compose config --quiet
docker compose up -d
docker compose ps
```
Stacks that attach to `proxy` require an existing Docker network of that name and
an appropriate reverse proxy. Published host ports still work independently of
Traefik. Check port conflicts before starting an alternative to a Kubernetes
service: DNS, STUN, and HTTP listeners can share the same host.
`docker compose down` keeps named volumes. Adding `-v` removes them.
## Preparing Kubernetes
The manifests assume Traefik CRDs, cert-manager, and a working storage provisioner.
PrometheusRule and ServiceMonitor resources also need the Prometheus Operator.
Replace the lab's hosts and addresses before using the configuration elsewhere.
Create a service's namespace, then prepare its ignored Secret from the example.
For example:
```sh
kubectl apply -f netbox/k8s/namespace.yaml
cp netbox/k8s/secrets.yaml.example netbox/k8s/secrets.yaml
$EDITOR netbox/k8s/secrets.yaml
kubectl apply -f netbox/k8s/secrets.yaml
```
The deploy workflow applies the tracked resources for marked services. Avoid
applying an entire `k8s/` directory blindly: some directories contain Helm values,
examples, and alternative routes. For a manual change, apply the selected manifest
explicitly and check the resulting rollout.
Shared database passwords must agree between the `database` namespace and each
application's Secret. Updating the PostgreSQL Secret does not change an existing
role's password; see the database README.
## Local checks
CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions:
```sh
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)"
export PATH="$tools_dir:$PATH"
ruff check .
ruff format --check .
actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
.gitea/workflows/sync-renovate-configmap.sh --check
```
The [workflow README](.gitea/README.md#checks) lists the rest of the checks.
Structure checks do not establish that local Secrets, mounted files, storage,
or external services are ready.
## Data and recovery
State lives outside Git: PVCs, Docker volumes, bind mounts, databases, and ignored
configuration. Keep backups of application data and the keys needed to read it.
An image rollback does not roll back database migrations or ConfigMap contents.
Many PVCs use the cluster's default StorageClass; monitoring explicitly uses
`local-path`. Check the PV reclaim policy before deleting a PVC or namespace.
The manifests do not provide a repository-wide backup schedule.
`incident-archive/` contains past incident notes. `.docs/storage-audit-instruction.md`
is a planning document, not evidence that NFS has been installed.
-22
View File
@@ -1,22 +0,0 @@
# AdGuard Home
DNS filtering with a web UI, DNS-over-TLS, and certificates from cert-manager.
The Kubernetes namespace is `adguard`. The workload uses `adguard-pvc` for
configuration and working data, and mounts the `adguard-certs` TLS Secret.
The LoadBalancer Service exposes DNS separately from the web ingress.
The Compose stack publishes TCP/UDP 53 and TCP 853 on the host. Prepare `conf/`
and `certs/` before starting it. Starting both DNS deployments on the same address
can cause a port conflict.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n adguard
kubectl get events -n adguard --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -12
View File
@@ -58,14 +58,10 @@ spec:
selector:
matchLabels:
app: adguard
strategy:
type: Recreate
template:
metadata:
labels:
app: adguard
annotations:
reloader.stakater.com/auto: "true"
spec:
containers:
- name: adguard
@@ -75,7 +71,7 @@ spec:
memory: "1.5Gi"
cpu: "300m"
requests:
memory: "512Mi"
memory: "500Mi"
cpu: "50m"
ports:
- containerPort: 3000
@@ -84,13 +80,6 @@ spec:
name: dns
- containerPort: 853
name: dot
readinessProbe:
tcpSocket:
port: dns
initialDelaySeconds: 5
periodSeconds: 5
successThreshold: 1
failureThreshold: 3
volumeMounts:
- name: adguard-data
mountPath: /opt/adguardhome/work
-12
View File
@@ -1,12 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: adguard-certs
namespace: adguard
spec:
secretName: adguard-certs
dnsNames:
- dns.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
+6 -5
View File
@@ -7,18 +7,21 @@ spec:
entryPoints:
- websecure
routes:
- match: Host(`dns.forust.xyz`)
- match: Host(`adguard.forust.xyz`) || Host(`dns.forust.xyz`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: adguard-service
port: 3000
- match: (Host(`dns.forust.xyz`)) && PathPrefix(`/dns-query`)
- match: (Host(`adguard.forust.xyz`) || Host(`dns.forust.xyz`)) && PathPrefix(`/dns-query`)
kind: Rule
services:
- name: adguard-service
port: 3000
tls:
secretName: adguard-certs
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -39,5 +42,3 @@ spec:
services:
- name: adguard-service
port: 3000
tls:
secretName: internal-wildcard-tls
-15
View File
@@ -1,15 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: adguard
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
-22
View File
@@ -1,22 +0,0 @@
# Authentik
Identity provider with separate server and worker deployments.
Kubernetes connects to the shared PostgreSQL service in `database`. Set
`AUTHENTIK_DB_PASSWORD` to the same value in both database and application Secrets.
Keep `AUTHENTIK_SECRET_KEY` with the backups.
Compose uses its own PostgreSQL 15 container and bind-mounted media and templates.
Its image defaults differ from Kubernetes; check both before an upgrade.
The worker mounts the Docker socket for Docker outpost management.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n authentik
kubectl get events -n authentik --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+5 -9
View File
@@ -34,8 +34,6 @@ spec:
selector:
matchLabels:
app: authentik-server
strategy:
type: Recreate
template:
metadata:
labels:
@@ -54,8 +52,8 @@ spec:
- containerPort: 9000
resources:
requests:
memory: "768Mi"
cpu: "100m"
memory: "700Mi"
cpu: "300m"
limits:
memory: "1.5Gi"
cpu: "1000m"
@@ -70,8 +68,6 @@ spec:
selector:
matchLabels:
app: authentik-worker
strategy:
type: Recreate
template:
metadata:
labels:
@@ -90,8 +86,8 @@ spec:
name: authentik-secrets
resources:
requests:
memory: "320Mi"
cpu: "100m"
memory: "512Mi"
cpu: "300m"
limits:
memory: "768Mi"
memory: "1Gi"
cpu: "700m"
-28
View File
@@ -1,28 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: authentik-prod-tls
namespace: authentik
spec:
secretName: authentik-prod-tls
dnsNames:
- auth.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: authentik
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+4 -3
View File
@@ -9,11 +9,14 @@ spec:
routes:
- match: Host(`auth.forust.xyz`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: authentik-server-service
port: 9000
tls:
secretName: authentik-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -29,5 +32,3 @@ spec:
services:
- name: authentik-server-service
port: 9000
tls:
secretName: internal-wildcard-tls
-18
View File
@@ -1,18 +0,0 @@
# cert-manager
Public ACME issuers and an internal certificate authority.
This directory contains chart values and issuer resources, not the controller
installation. Install the cert-manager chart with CRDs and the settings in
`k8s/cert-manager-values.yaml` before applying the issuers.
`clusterissuer.yaml` defines staging and production Let's Encrypt issuers.
They use HTTP-01 through the Traefik ingress class. Public DNS and inbound HTTP
reachability must work for the requested names before issuance.
`internal-ca.yaml` bootstraps the internal CA. Keep its private-key Secret backed
up; the tracked `.crt` is only a public certificate.
This directory has no `k8s/active` marker. Apply the issuer files deliberately;
`kubectl apply` does not interpret the Helm values file.
See the [repository README](../README.md) for deployment selection.
@@ -1,9 +0,0 @@
crds:
enabled: true
prometheus:
servicemonitor:
enabled: true
interval: 60s
scrapeTimeout: 30s
labels:
release: prometheus-stack
-29
View File
@@ -1,29 +0,0 @@
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-staging
spec:
acme:
email: bobrovod@national.shitposting.agency
server: https://acme-staging-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: letsencrypt-staging-account-key
solvers:
- http01:
ingress:
class: traefik
---
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
email: bobrovod@national.shitposting.agency
server: https://acme-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: letsencrypt-prod-account-key
solvers:
- http01:
ingress:
class: traefik
@@ -1,30 +0,0 @@
-----BEGIN CERTIFICATE-----
MIIFFjCCAv6gAwIBAgIUetKpTfEDOn2985FFMu6G26itT+wwDQYJKoZIhvcNAQEN
BQAwIzEhMB8GA1UEAxMYaG9tZWxhYiBpbnRlcm5hbCByb290IENBMB4XDTI2MDky
MzEyNDA0N1oXDTM2MDkyMDEyNDA0N1owIzEhMB8GA1UEAxMYaG9tZWxhYiBpbnRl
cm5hbCByb290IENBMIICIjANBgkqhkiG9w0BAQEFAAOCAg8AMIICCgKCAgEAvmNP
ZCOoD8NtNuYJKVXBlTPjX7D7sJCSK5neH7ZbYV5+lmUlEErY8Mik7j37V5k5NfpF
Ig85pOjP7RckTPz5V6ek3yaN40s4AL053sN5ZPauDVYjalaEHTgj5sEMqlLACQWI
yZmJOZspZykae8dIpQnqCoFpRT4FurJ78v4a0ylnFVLMQn/lyCHedwTjkEdtYWYr
ccJy8vQwqkzs/rWvEH1lDqZhennLOrmcCjfonG7D/pruMn4z+6E28p4+ejkRrI6x
luak3KnpT1XMeHtgU21hiRGaMDBchHMFgAhnY1qosymKenXvfTZItwgjZbwa1hJI
GAiDm+jQDKMjzRZ3rH6Xfc0auUcykNz73PpNu1NGm78nndXwCXcXn1LFKNQJ+r1U
sJiyAmUZmXVn4aM4OMf2F38k7wTYIKg7nRGaUkNeKDlNkjA4HvgWw+jwO1KmdHQ/
mOem1rosDWHRK01wg+Gga9mQCnhNhxglg3t/UeSic6uOaRsvaz4qkzHq8MbCujVz
DpKQjqdikYOAXZOs4KlBLWrS7NaK4NzfSD02pBUErh54ruJfY/bWz9KyXzBD/lQZ
VUTKyvUVB0bkVHEdf1jJmX3H4IZRQSF5JPqOBotW6bJI5fEGNBvj9Zxy4nm2WWGz
yyP3uWsQz8U/Wdx9nXZLHInTZBsvgLYtKUAWA30CAwEAAaNCMEAwDgYDVR0PAQH/
BAQDAgKkMA8GA1UdEwEB/wQFMAMBAf8wHQYDVR0OBBYEFEkKm2rxPaK6+O9WD80z
BLC6F9QsMA0GCSqGSIb3DQEBDQUAA4ICAQAnFyHz97Umf5VIu+dKTJid7C73VugJ
TIar/xJBs/4CxP+znBxhJjXygRoyIfzoVGWcB2ZSL//vL78Qlts79K/Imc9a4RFF
wMvCxsRXAEQ4TpeWi3ophPNcs4rhsP+gQKQFtnyKP9519bqpfxp0bTqwOV2o18fn
za7rlQViiEnNV58j7CVoM9+mJvVVfBEX1Km+GyJL9GadzbIQ7FxClVJZefCbft93
zHVk9gDOw8ys1XGSR2OUCyCLinXO6mqS16CmBb2MAKXq/YyH7E0N8iotAPGtfA8V
M/0ddy947rY0xCrtECfWwvGQpJS7NRv/Z9b2jCfXrI5LXmL2nfQRg0y9GE4Vjwr+
WxtGU5jOeFt0jQ+xRzcgG0Op+qK3x55l5LSo2hOcOVYbiHxcHEJFgwNi1ADeBFwb
q/HdysfURSOghqjIpMMAUabBp+DBUg2EUF7pIaUqbdqExFYcr9EYisEMiNsmKmN+
8ZbcOeerFKDQj+t/R0bFXa7UBn2UWsjI8zlR74aa2kLDXwtyz/XlO/FlYm66eBFo
2/eYUSeU+S4ej+wUAs/dvjF7f190/DUQGuwOTlLTahqWDztmhCk7qzbECu56CwKT
E5Ect2P72UleYwdblkVOVd352AmiwEzdOaziIRrPh8uenEknH6JBYPO3mjk7cCg+
GPGcNIBctjdXhg==
-----END CERTIFICATE-----
-33
View File
@@ -1,33 +0,0 @@
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: selfsigned
spec:
selfSigned: {}
---
# Homelab internal root CA (10y). Install the .crt on clients (see below).
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-ca-root
namespace: cert-manager
spec:
isCA: true
commonName: homelab internal root CA
duration: 87600h
renewBefore: 7200h
secretName: internal-ca-root
privateKey:
algorithm: RSA
size: 4096
issuerRef:
name: selfsigned
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: internal-ca
spec:
ca:
secretName: internal-ca-root
-4
View File
@@ -1,4 +0,0 @@
apiVersion: v1
kind: Namespace
metadata:
name: cert-manager
-22
View File
@@ -1,22 +0,0 @@
# Cloudflare DDNS
Updates the lab DNS records when the public address changes.
Kubernetes runs in `default` with host networking and reads `cfddns-secrets`.
The Compose stack also uses host networking. Configure the API token and domain
list from the relevant example; keep DNS names consistent with the ingress rules.
`config.json.example` is a separate configuration example. The current Compose
file does not mount a config.json file. Check configuration against the pinned
DDNS image when changing between environment and file-based settings.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n default
kubectl get events -n default --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+2 -4
View File
@@ -9,8 +9,6 @@ spec:
selector:
matchLabels:
app: cfddns
strategy:
type: Recreate
template:
metadata:
labels:
@@ -24,10 +22,10 @@ spec:
imagePullPolicy: Always
resources:
requests:
memory: "32Mi"
memory: "20Mi"
cpu: "30m"
limits:
memory: "128Mi"
memory: "64Mi"
cpu: "50m"
envFrom:
- secretRef:
-21
View File
@@ -1,21 +0,0 @@
# Checkmk
Checkmk Raw monitoring site with web and agent-receiver ingress.
The site data lives in `checkmk-sites-pvc` on Kubernetes and the `sites` named
volume on Compose. The agent receiver has a separate TCP route; enabling the
web route alone does not expose it.
Prepare the password in the service env or Secret example. Inspect the Checkmk
container logs during the first site creation.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n checkmk
kubectl get events -n checkmk --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
File renamed without changes.
-28
View File
@@ -1,28 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: checkmk-prod-tls
namespace: checkmk
spec:
secretName: checkmk-prod-tls
dnsNames:
- cmk.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: checkmk
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
-2
View File
@@ -24,8 +24,6 @@ spec:
selector:
matchLabels:
app: checkmk
strategy:
type: Recreate
template:
metadata:
labels:
+4 -3
View File
@@ -9,11 +9,14 @@ spec:
routes:
- match: Host(`cmk.forust.xyz`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: checkmk-service
port: 5000
tls:
secretName: checkmk-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRouteTCP
@@ -45,5 +48,3 @@ spec:
services:
- name: checkmk-service
port: 5000
tls:
secretName: internal-wildcard-tls
-21
View File
@@ -1,21 +0,0 @@
# Cloudflare Tunnel
A Kubernetes connector for an existing Cloudflare tunnel.
The Deployment runs in `default` and reads its token from the ignored Secret
created from `k8s/secret.yaml.example`. Create the tunnel and its hostname rules
in Cloudflare before starting the connector.
There is no Compose file or `k8s/active` marker. Apply the Secret first, then
`k8s/deployment.yaml` when this tunnel is needed.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n default
kubectl get events -n default --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
File renamed without changes.
+3 -5
View File
@@ -9,8 +9,6 @@ spec:
selector:
matchLabels:
app: cloudflared
strategy:
type: Recreate
template:
metadata:
labels:
@@ -18,7 +16,7 @@ spec:
spec:
containers:
- name: cloudflared
image: cloudflare/cloudflared:2026.9.3
image: cloudflare/cloudflared:2026.9.1
imagePullPolicy: IfNotPresent
args:
- tunnel
@@ -32,8 +30,8 @@ spec:
key: TUNNEL_TOKEN
resources:
requests:
memory: "128Mi"
memory: "32Mi"
cpu: "30m"
limits:
memory: "256Mi"
memory: "128Mi"
cpu: "200m"
-22
View File
@@ -1,22 +0,0 @@
# File converters
ConvertX for server-side conversion and BentoPDF for PDF tools.
ConvertX persists files in `convertx-pvc`; BentoPDF has no persistent volume.
Kubernetes configuration includes a local `config.yaml.example`, excluded from
normal deployment. Copy and apply the real ConfigMap separately where required.
Compose publishes ConvertX on host port 9992 as well as attaching it to the
proxy network. Replace the authentication settings from `.env.example` before
exposing it outside the lab.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n converters
kubectl get events -n converters --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+2 -2
View File
@@ -1,7 +1,7 @@
services:
convertx:
container_name: convertx
image: ghcr.io/c4illin/convertx:v0.19.0
image: ghcr.io/c4illin/convertx:v0.18.0
restart: unless-stopped
ports:
- "9992:3000"
@@ -54,7 +54,7 @@ services:
- "traefik.http.routers.bentopdf.tls.certresolver=letsencrypt"
- "traefik.http.routers.bentopdf.tls=true"
# Local router
- "traefik.http.routers.bentopdf-local.rule=Host(`pdf.workstation.internal`)"
- "traefik.http.routers.bentopdf-local.rule=Host(`pdf.wokstation.internal`)"
- "traefik.http.routers.bentopdf-local.entrypoints=websecure"
- "traefik.http.routers.bentopdf-local.tls=true"
# Dev router
+2 -3
View File
@@ -31,13 +31,12 @@ spec:
name: bentopdf
ports:
- containerPort: 8080
# p95 4M, max 11M over 7 days. Was 50Mi/700Mi.
resources:
requests:
memory: "32Mi"
memory: "50Mi"
cpu: "50m"
ephemeral-storage: "100Mi"
limits:
memory: "128Mi"
memory: "700Mi"
cpu: "700m"
ephemeral-storage: "5Gi"
-42
View File
@@ -1,42 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: convertx-prod-tls
namespace: converters
spec:
secretName: convertx-prod-tls
dnsNames:
- forust.xyz
- www.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: bentopdf-prod-tls
namespace: converters
spec:
secretName: bentopdf-prod-tls
dnsNames:
- pdf.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: converters
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -6
View File
@@ -20,15 +20,13 @@ spec:
selector:
matchLabels:
app: convertx
strategy:
type: Recreate
template:
metadata:
labels:
app: convertx
spec:
containers:
- image: ghcr.io/c4illin/convertx:v0.19.0
- image: ghcr.io/c4illin/convertx:v0.18.0
name: convertx
envFrom:
- configMapRef:
@@ -40,14 +38,13 @@ spec:
volumeMounts:
- mountPath: /data
name: data
# p95 85M, max 136M over 7 days, spikes while converting. Was 250Mi/1.5Gi.
resources:
requests:
memory: "128Mi"
memory: "250Mi"
cpu: "100m"
limits:
cpu: "1500m"
memory: "512Mi"
memory: "1.5Gi"
volumes:
- name: data
persistentVolumeClaim:
+2 -7
View File
@@ -14,7 +14,7 @@ spec:
- name: convertx-service
port: 3000
tls:
secretName: convertx-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -31,9 +31,6 @@ spec:
services:
- name: convertx-service
port: 3000
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -50,7 +47,7 @@ spec:
- name: bentopdf-service
port: 8080
tls:
secretName: bentopdf-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -66,5 +63,3 @@ spec:
services:
- name: bentopdf-service
port: 8080
tls:
secretName: internal-wildcard-tls
-25
View File
@@ -1,25 +0,0 @@
# CrowdSec
Helm values, dashboards, network policy, and a maintenance CronJob.
Install CrowdSec separately using `k8s/crowdsec-values.yaml`; the deploy
workflow does not have a CrowdSec Helm release entry. There is no `k8s/active`
marker in this directory.
The LAPI policy and janitor run in `crowdsec`. The dashboard ConfigMaps are in
`prometheus` for Grafana's sidecar. The janitor has its own ServiceAccount and
namespace Role. Review its script and schedule before enabling cleanup.
Traefik's values state that enforcement moved to a host firewall bouncer. This
repository does not install that host component.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n crowdsec
kubectl get events -n crowdsec --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+14
View File
@@ -0,0 +1,14 @@
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: crowdsec-bouncer
namespace: crowdsec
spec:
plugin:
crowdsec-bouncer:
enabled: true
LogLevel: INFO
CrowdsecMode: live
CrowdsecLapiScheme: http
CrowdsecLapiHost: crowdsec-service.crowdsec.svc.cluster.local:8080
CrowdsecLapiKeyFile: "/etc/traefik/secrets/traefik-api-key"
+4 -105
View File
@@ -16,12 +16,6 @@ agent:
value: crowdsecurity/traefik crowdsecurity/base-http-scenarios
- name: DISABLE_COLLECTIONS
value: crowdsecurity/sshd
# Bans on 401/403 bursts hurt more than they protect: with L3 enforcement
# a false positive cuts the IP off everything (SSH included), and past
# incidents show legit automation (deploy runner, mesh peers, registry
# pulls) tripping this probe. Probing/XSS/SQLi/CVE scenarios stay.
- name: DISABLE_SCENARIOS
value: crowdsecurity/http-generic-bf
metrics:
enabled: true
serviceMonitor:
@@ -64,44 +58,9 @@ config:
reason: "Mobile IP whitelist"
cidr:
- "84.245.64.0/18"
# CrowdSec's own guidance: CIDR allowlisting belongs at the parser stage.
# A parser whitelist discards the event before it reaches a bucket, so
# these addresses never produce an overflow and never become a decision.
# A postoverflow whitelist is checked only *after* the ban exists, and
# the bouncer answers 403 for as long as it does - which is a window we
# do not want the deploy sitting in.
local-network.yaml: |
name: forust/local-network
description: "Whitelist loopback, private and VPN networks"
whitelist:
reason: "Local network"
cidr:
- "127.0.0.0/8"
- "10.0.0.0/8"
- "172.16.0.0/12"
- "192.168.0.0/16"
# CGNAT range (RFC 6598). The workstation and the k0s node live
# here on WireGuard, and 100.64.0.0/10 is not covered by the
# RFC 1918 blocks above.
- "100.64.0.0/10"
- "169.254.0.0/16"
- "fc00::/7"
- "fe80::/10"
vps-whitelist.yaml: |
name: forust/vps-whitelist
description: "Whitelist static VPS"
whitelist:
reason: "VPS"
ip:
- "193.181.211.79"
postoverflows:
s01-whitelist:
# The one whitelist that has to stay here: resolving a hostname is a
# network call, and the docs put expensive lookups in postoverflows on
# purpose - it runs only when a bucket actually overflows.
# ddns.forust.xyz is the public home address, not a private one, so
# forust/local-network does not cover it.
home-dynamic-ip.yaml: |
name: forust/home-dynamic-ip
description: "Whitelist home dynamic IP"
@@ -110,59 +69,6 @@ config:
expression:
- evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz")
# LAPI-only main config override, merged over config.yaml. NOTE: the
# chart's own default for this key is REPLACED, not merged, so its
# auto_registration block is repeated verbatim below - drop it and the
# agent can no longer register itself.
config.yaml.local: |
api:
server:
auto_registration: # Activate if not using TLS for authentication
enabled: true
token: "${REGISTRATION_TOKEN}" # /!\ Do not modify this variable (auto-generated and handled by the chart)
allowed_ranges: # /!\ Make sure to adapt to the pod IP ranges used by your cluster
- "127.0.0.1/32"
- "192.168.0.0/16"
- "10.0.0.0/8"
- "172.16.0.0/12"
# This homelab has no egress to console.crowdsec.cloud: DNS does
# not resolve. The LAPI kept trying anyway ("Signal push: N
# signals to push", "capi metrics: sending" every 10s) and each
# attempt sat on a resolver timeout WHILE HOLDING A WRITE
# TRANSACTION, which is what kept stalling per-request decision
# lookups even with WAL enabled. Nothing to share and nothing to
# pull - turn the Central API off instead of letting it block the
# only database writer we have.
online_client:
sharing: false
pull:
community: false
blocklists: false
disable_usage_metrics_export: true
db_config:
# SQLite without WAL serialises every reader behind the writer's
# rollback journal, and the LAPI writes constantly: the agent pushes
# Traefik alerts read from Loki, the metrics collector counts
# decisions, the bouncer touches "last pull" on every request.
# Symptom: decision lookups taking 10-30s (and a second connection
# that could not even open the database) while the LAPI sat at 28m
# CPU - the process was blocked in fsync, not computing. Every
# bouncer-protected request then blew through the plugin timeout and
# fail-closed with 403, on every site at once.
# The PVC is local-path-retain (hostPath), not a network share, so
# WAL is safe here; the crowdsec docs recommend it for exactly this
# ("allowing more concurrency in SQLite that will improve
# performances in most scenarios").
use_wal: true
# Keeps the alert table bounded. At the 5000/7d default the file
# reached 54MB in 15 days off the Traefik access log alone, and the
# metrics collector counts decisions on a timer; a smaller working
# set means fewer full scans. Crowdsec only prunes - SQLite never
# shrinks the file, so the size stays until a manual VACUUM.
flush:
max_items: 1000
max_age: 24h
lapi:
env:
- name: COLLECTIONS
@@ -184,20 +90,13 @@ lapi:
enabled: true
size: 1Gi
storageClassName: local-path-retain
# LAPI answers a blocking /v1/decisions lookup for EVERY bouncer-protected
# request (whole Traefik front door), so it is the hot path of the proxy.
# At 400m/500Mi it went CPU-throttled and idle lookups measured 1.3-7.4s,
# which pushed requests into the bouncer's fail-closed 403.
# Single replica on purpose: LAPI is stateful (BoltDB on the `data` PVC,
# credentials on the `config` PVC) - two replicas sharing those RWO
# volumes would corrupt the decision store. Scale up CPU, not replicas.
resources:
limits:
cpu: 1500m
memory: 1Gi
requests:
cpu: 250m
cpu: 400m
memory: 500Mi
requests:
cpu: 50m
memory: 150Mi
service:
type: ClusterIP
storeLAPICscliCredentialsInSecret: true
-7
View File
@@ -32,13 +32,6 @@
# on their own - same name + same password);
# 4. prune bouncer entries idle for 30d.
#
# It used to also delete LePresidente/http-generic-403-bf decisions hourly.
# That was a workaround for the bouncer failing closed on a slow LAPI and
# 403-ing the deploy runner into a 4h ban. The bouncer now polls decisions
# into a cache and never blocks on an unreachable LAPI, so it cannot
# manufacture those 403s any more, and the scenario only fires against real
# scanners - deleting their decisions hourly was undoing a working ban.
#
# Manual apply (crowdsec/k8s is NOT managed by deploy.yaml):
# kubectl apply -f crowdsec/k8s/janitor-cronjob.yaml
# Force a run:
-21
View File
@@ -1,21 +0,0 @@
# Dockmon
Docker management UI that talks to the host Docker daemon.
Both runtimes mount `/var/run/docker.sock`. On Kubernetes the socket belongs
to the node hosting the pod, so this is not a cluster-wide container manager.
Compose stores application data in a named volume. Kubernetes uses a StatefulSet
with a volume claim template. Its ServersTransport is specific to the upstream
connection; keep it with the ingress resources.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n dockmon
kubectl get events -n dockmon --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,6 +1,6 @@
services:
dockmon:
image: darthnorse/dockmon:2.5.0
image: darthnorse/dockmon:2.4.5
container_name: dockmon
restart: unless-stopped
# ports:
File renamed without changes.
-28
View File
@@ -1,28 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: dockmon-prod-tls
namespace: dockmon
spec:
secretName: dockmon-prod-tls
dnsNames:
- dockmon.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: dockmon
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+1 -1
View File
@@ -29,7 +29,7 @@ spec:
spec:
containers:
- name: dockmon
image: darthnorse/dockmon:2.5.0
image: darthnorse/dockmon:2.4.5
ports:
- containerPort: 443
volumeMounts:
+3 -3
View File
@@ -18,13 +18,15 @@ spec:
- match: Host(`dockmon.forust.xyz`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
- name: security-headers@file
services:
- name: dockmon-service
port: 443
serversTransport: dockmon-transport
tls:
secretName: dockmon-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -41,5 +43,3 @@ spec:
- name: dockmon-service
port: 443
serversTransport: dockmon-transport
tls:
secretName: internal-wildcard-tls
-151
View File
@@ -1,151 +0,0 @@
# Repository review
Reviewed the tracked tree at `cc9c3de` and read the live workstation state on
6 October 2026. Changes are split into documentation and individual fix branches,
all based on that main commit. The original local checkout and its uncommitted
monitoring changes were preserved. No deployment was performed.
## Confirmed problems with prepared fixes
| Priority | Problem and consequence | Fix branch |
| -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------- |
| High | `APPLY_PRUNE=true` is passed to each individual manifest apply. Each invocation sees only that file's desired objects and can delete other resources selected by the shared label. | `fix/deploy-prune-guard` |
| High | Deploy validates Compose with interpolation and env/path resolution disabled. Required settings can pass validation and then fail during apply after other workloads have changed. | `fix/deploy-validation` |
| Medium | Secret validation is text-based and compares names across all namespaces. A Secret elsewhere can hide a missing local Secret; mounted Secrets are also missed. | `fix/deploy-validation` |
| Medium | Compose CI misses `postgres/shared-compose.yaml`, `netbird/client.compose.yaml`, and `renovate/renovate-compose.yaml`. | `fix/deploy-validation` |
| Medium | NetBird Compose mounts `entrypoint.sh`, but it is absent. Its README also calls a missing `setup.sh`; a fresh checkout cannot start this stack as documented. | `fix/netbird-compose-runtime` |
| Medium | Glance's CSS mount uses `glance-config`, whose keys do not include `user.css`. That key is in `glance-assets`; the pod's subPath mount cannot be prepared correctly. | `fix/glance-assets` |
| Medium | The shared PostgreSQL initializer requires `NETBOX_DB_PASSWORD`, but the Compose env example omits it. Following the example leaves first initialization incomplete. | `fix/postgres-env-example` |
| Medium | EDU's Compose env example uses old credential names and full URL variables, while the code reads `KEEPER_*` and paths under `EDU_URL_BASE`. | `fix/session-keeper-reliability` |
| Medium | Session keeper HTTP calls have no timeouts. Its Redis cookie never expires, probes only check existence, and its logs include cookies. A hung or failed refresh can leave a stale session appearing ready. | `fix/session-keeper-reliability` |
| Medium | AdGuard's DoH and SearXNG's Compose rules put Boolean expressions inside `Host(...)`. They are invalid router expressions despite valid YAML. | `fix/compose-router-rules` |
Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is
described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/).
The fix retains the DoH path constraint for both hostnames.
The prune fix deliberately rejects the unsafe option. It does not introduce
automatic deletion under a different implementation. Prune defaults to false,
and no tracked resource currently carries the selector label, so this is a
latent defect rather than evidence of a live deletion incident.
The session fix bounds HTTP and Redis calls, validates required credentials,
sets a cookie lifetime of two refresh intervals, and marks success only after
publishing the verified cookie. With the default ten-minute interval, an outage
longer than twenty minutes will make the existing Redis-key readiness checks fail.
That is an intentional change from indefinite apparent readiness.
The deployment fix extracts required pod Secret references from rendered JSON,
checks their namespaces, includes init containers, image-pull credentials, and
mounted/projected Secrets, and honors optional references. Ingress TLS Secrets
issued by cert-manager are not treated as pre-existing pod prerequisites.
It checks existence/access, not every key's contents or application validity.
## Live workstation observations
The SSH alias `workstation` is reachable. It has one Ready control-plane node,
Kubernetes `v1.35.4+k0s`, and a Docker daemon alongside containerd. At inspection,
no pods were Pending or in another non-running, non-completed phase. This is a
point-in-time observation, not a complete application health test.
The deployment checkout at `/srv/homelab` is on main commit `2adf17c`, behind the
reviewed local commit. It has untracked host configuration and a separate
`userbot/` directory. It was not reset or cleaned.
| Observed difference | Implication |
| ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| VictoriaMetrics and vmalert are running; the Prometheus StatefulSet has zero replicas. | A monitoring migration is already in progress outside committed main. Deploying the old Helm values can overwrite those settings. |
| Homarr, Cloudflared, and Reloader are installed without their current Git active markers. | Installed services and marker-selected services are different inventories. Missing markers do not establish that a service is stopped. |
| Cloudflare DDNS is running in both Docker and Kubernetes. | Confirm which instance should own DNS updates and whether their domain lists overlap before retiring either one. Secret values were not inspected. |
| Traefik's LoadBalancer exposes port 8080 at `192.168.80.2`. | The direct API listener is deployed; its external reachability was not tested. |
| Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. |
| NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. |
The monitoring files already modified in the user's local tree correspond to the
live migration. They are excluded from these branches. Reconcile that work before
using this review's baseline to deploy monitoring.
## Remaining work
These need recovery design or infrastructure decisions rather than a small
configuration correction:
- **SSH apply retries can replace the rollback baseline.** `ssh-run.sh` retries
exit 255, including `apply-k8s`; every new invocation publishes a fresh snapshot.
If the first attempt already changed workloads, the retry snapshots that partial
state. Preserve a run-specific original baseline and verify it across retries.
- **Rollback can exceed the job budget.** Verification is parallel, but
`rollback_workloads` is serial with a five-minute limit per workload. The
thirty-minute job budget can expire before recovery finishes. Bound recovery
concurrency and account for both phases before choosing a new timeout.
- **Snapshot collection is allowed to fail.** Generation and workload snapshot
errors are warnings; verify can fall back to all workloads. A snapshot failure
must not permit unrelated workloads to be selected for automatic undo.
- **Rollback uses the previous revision, not the captured revision.** `rollout undo`
without an explicit revision cannot guarantee restoration to the snapshot after
retries or intervening rollouts. First deployments also have no previous revision.
- **Manual deploy dispatch bypasses the CI-success trigger.** Either validate the
target commit's successful CI run or document manual dispatch as an operator
override with its own required checks.
- **Direct Traefik API exposure is unauthenticated.** The latest local commit
explicitly added it for Homarr. Preserve that integration while choosing a
cluster-internal authenticated path or a verified network restriction; do not
simply disable an integration that is already in use.
- **Storage retention and backup are not reproducible as a whole.** Defaults and
several important PV policies are Delete. There is no repository-wide backup
schedule. Existing PVC StorageClass changes require migration rather than an
in-place YAML edit.
- **MeTube downloads are temporary on Kubernetes.** `/downloads` is a 20 GiB
emptyDir. Decide whether pod replacement should discard files or whether it
should use persistent storage. Compose uses a host directory instead.
- **First-time activation needs a bootstrap path.** Deploy validation dry-runs
namespaced resources before the apply stage creates namespaces and installs
selected charts. On a fresh cluster, missing namespaces and CRDs need separate
preparation; activation is not a complete installer.
## Validation
Baseline lint checks passed for Python, shell, workflows, YAML, standard Compose
files, and Kubernetes resources with available schemas. Kubeconform found 347
resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources.
That skip count matters: passing schema validation does not validate Traefik rule
strings or other controller-specific behavior.
Fix validation covers:
- Compose discovery of manual entry points, rejection of required-variable gaps,
namespace-scoped and optional Secret references, and API/render failures.
- NetBird setup idempotence, preservation of existing keys, file permissions,
runtime rendering, and rejection of invalid trusted proxy CIDRs.
- Session refresh success and failure paths, timeouts, cookie expiry, log redaction,
missing credentials, and nonpositive refresh intervals.
- Correct Glance ConfigMap key selection and PostgreSQL initializer/env alignment.
- YAML and Compose structure for the corrected router rules, compared with the
documented Traefik grammar. They were not exercised on the live proxy.
- Prune rejection before any cluster invocation.
All seven fix branches and the documentation branch merged together without
conflicts in a disposable validation worktree. The combined tree passed the
CI-equivalent local checks, Markdown formatting/lint and link checks, all 35
Compose structure checks, and 11 Python regression tests plus the shell
validation regressions. CRD server-side validation and live rollout tests were
not run.
Runtime tests use fixtures and mocks, not production credentials. Live checks read
workload metadata, storage policies, chart versions, and container state only.
They did not read Secret contents or change services.
## Reloader follow-up
`fix/reloader-integration` adds the active marker and opt-in annotations to 28
application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets.
It corrects AdGuard's misplaced pod-template annotation. The Helm settings use
annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and
CronJobs. PostgreSQL workloads are excluded because their credential variables
and init scripts are only effective on an empty data directory.
The controller was already running on workstation when inspected. Its live
configuration is unchanged by the branch: merge and deploy the integration to
apply the new policy and application annotations. Configuration reload behavior
was checked against the pinned chart, with Helm rendering and manifest validation;
no production configuration was changed to provoke a test restart.
-20
View File
@@ -1,20 +0,0 @@
# Downtify
Download UI with a persistent downloads directory.
Compose stores downloads under `Downtify_downloads/`; Kubernetes uses
`downtify-downloads-pvc`. The ingress manifests reference shared infrastructure,
so check certificate and middleware availability before enabling them.
Back up downloads separately if they need to survive storage replacement.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n downtify
kubectl get events -n downtify --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,7 +1,7 @@
services:
downtify:
container_name: downtify
image: ghcr.io/henriquesebastiao/downtify:3.4.0
image: ghcr.io/henriquesebastiao/downtify:2.13.0
restart: unless-stopped
# ports:
# - '7077:8000'
+1 -3
View File
@@ -20,8 +20,6 @@ spec:
selector:
matchLabels:
app: downtify
strategy:
type: Recreate
template:
metadata:
labels:
@@ -29,7 +27,7 @@ spec:
spec:
containers:
- name: downtify
image: ghcr.io/henriquesebastiao/downtify:3.4.0
image: ghcr.io/henriquesebastiao/downtify:2.13.0
ports:
- containerPort: 8000
volumeMounts:
+3 -3
View File
@@ -10,12 +10,14 @@ spec:
- match: Host(`downtify.forust.xyz`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
- name: security-chain@file
services:
- name: downtify-service
port: 8000
tls:
secretName: downtify-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -31,5 +33,3 @@ spec:
services:
- name: downtify-service
port: 8000
tls:
secretName: internal-wildcard-tls
+13
View File
@@ -0,0 +1,13 @@
FROM python:3.9-alpine
WORKDIR /app
# Установка зависимостей
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Копирование кода
COPY main.py .
COPY .env .
# Запуск бота
CMD ["python", "-u", "main.py"]
+373
View File
@@ -0,0 +1,373 @@
Mozilla Public License Version 2.0
==================================
1. Definitions
--------------
1.1. "Contributor"
means each individual or legal entity that creates, contributes to
the creation of, or owns Covered Software.
1.2. "Contributor Version"
means the combination of the Contributions of others (if any) used
by a Contributor and that particular Contributor's Contribution.
1.3. "Contribution"
means Covered Software of a particular Contributor.
1.4. "Covered Software"
means Source Code Form to which the initial Contributor has attached
the notice in Exhibit A, the Executable Form of such Source Code
Form, and Modifications of such Source Code Form, in each case
including portions thereof.
1.5. "Incompatible With Secondary Licenses"
means
(a) that the initial Contributor has attached the notice described
in Exhibit B to the Covered Software; or
(b) that the Covered Software was made available under the terms of
version 1.1 or earlier of the License, but not also under the
terms of a Secondary License.
1.6. "Executable Form"
means any form of the work other than Source Code Form.
1.7. "Larger Work"
means a work that combines Covered Software with other material, in
a separate file or files, that is not Covered Software.
1.8. "License"
means this document.
1.9. "Licensable"
means having the right to grant, to the maximum extent possible,
whether at the time of the initial grant or subsequently, any and
all of the rights conveyed by this License.
1.10. "Modifications"
means any of the following:
(a) any file in Source Code Form that results from an addition to,
deletion from, or modification of the contents of Covered
Software; or
(b) any new file in Source Code Form that contains any Covered
Software.
1.11. "Patent Claims" of a Contributor
means any patent claim(s), including without limitation, method,
process, and apparatus claims, in any patent Licensable by such
Contributor that would be infringed, but for the grant of the
License, by the making, using, selling, offering for sale, having
made, import, or transfer of either its Contributions or its
Contributor Version.
1.12. "Secondary License"
means either the GNU General Public License, Version 2.0, the GNU
Lesser General Public License, Version 2.1, the GNU Affero General
Public License, Version 3.0, or any later versions of those
licenses.
1.13. "Source Code Form"
means the form of the work preferred for making modifications.
1.14. "You" (or "Your")
means an individual or a legal entity exercising rights under this
License. For legal entities, "You" includes any entity that
controls, is controlled by, or is under common control with You. For
purposes of this definition, "control" means (a) the power, direct
or indirect, to cause the direction or management of such entity,
whether by contract or otherwise, or (b) ownership of more than
fifty percent (50%) of the outstanding shares or beneficial
ownership of such entity.
2. License Grants and Conditions
--------------------------------
2.1. Grants
Each Contributor hereby grants You a world-wide, royalty-free,
non-exclusive license:
(a) under intellectual property rights (other than patent or trademark)
Licensable by such Contributor to use, reproduce, make available,
modify, display, perform, distribute, and otherwise exploit its
Contributions, either on an unmodified basis, with Modifications, or
as part of a Larger Work; and
(b) under Patent Claims of such Contributor to make, use, sell, offer
for sale, have made, import, and otherwise transfer either its
Contributions or its Contributor Version.
2.2. Effective Date
The licenses granted in Section 2.1 with respect to any Contribution
become effective for each Contribution on the date the Contributor first
distributes such Contribution.
2.3. Limitations on Grant Scope
The licenses granted in this Section 2 are the only rights granted under
this License. No additional rights or licenses will be implied from the
distribution or licensing of Covered Software under this License.
Notwithstanding Section 2.1(b) above, no patent license is granted by a
Contributor:
(a) for any code that a Contributor has removed from Covered Software;
or
(b) for infringements caused by: (i) Your and any other third party's
modifications of Covered Software, or (ii) the combination of its
Contributions with other software (except as part of its Contributor
Version); or
(c) under Patent Claims infringed by Covered Software in the absence of
its Contributions.
This License does not grant any rights in the trademarks, service marks,
or logos of any Contributor (except as may be necessary to comply with
the notice requirements in Section 3.4).
2.4. Subsequent Licenses
No Contributor makes additional grants as a result of Your choice to
distribute the Covered Software under a subsequent version of this
License (see Section 10.2) or under the terms of a Secondary License (if
permitted under the terms of Section 3.3).
2.5. Representation
Each Contributor represents that the Contributor believes its
Contributions are its original creation(s) or it has sufficient rights
to grant the rights to its Contributions conveyed by this License.
2.6. Fair Use
This License is not intended to limit any rights You have under
applicable copyright doctrines of fair use, fair dealing, or other
equivalents.
2.7. Conditions
Sections 3.1, 3.2, 3.3, and 3.4 are conditions of the licenses granted
in Section 2.1.
3. Responsibilities
-------------------
3.1. Distribution of Source Form
All distribution of Covered Software in Source Code Form, including any
Modifications that You create or to which You contribute, must be under
the terms of this License. You must inform recipients that the Source
Code Form of the Covered Software is governed by the terms of this
License, and how they can obtain a copy of this License. You may not
attempt to alter or restrict the recipients' rights in the Source Code
Form.
3.2. Distribution of Executable Form
If You distribute Covered Software in Executable Form then:
(a) such Covered Software must also be made available in Source Code
Form, as described in Section 3.1, and You must inform recipients of
the Executable Form how they can obtain a copy of such Source Code
Form by reasonable means in a timely manner, at a charge no more
than the cost of distribution to the recipient; and
(b) You may distribute such Executable Form under the terms of this
License, or sublicense it under different terms, provided that the
license for the Executable Form does not attempt to limit or alter
the recipients' rights in the Source Code Form under this License.
3.3. Distribution of a Larger Work
You may create and distribute a Larger Work under terms of Your choice,
provided that You also comply with the requirements of this License for
the Covered Software. If the Larger Work is a combination of Covered
Software with a work governed by one or more Secondary Licenses, and the
Covered Software is not Incompatible With Secondary Licenses, this
License permits You to additionally distribute such Covered Software
under the terms of such Secondary License(s), so that the recipient of
the Larger Work may, at their option, further distribute the Covered
Software under the terms of either this License or such Secondary
License(s).
3.4. Notices
You may not remove or alter the substance of any license notices
(including copyright notices, patent notices, disclaimers of warranty,
or limitations of liability) contained within the Source Code Form of
the Covered Software, except that You may alter any license notices to
the extent required to remedy known factual inaccuracies.
3.5. Application of Additional Terms
You may choose to offer, and to charge a fee for, warranty, support,
indemnity or liability obligations to one or more recipients of Covered
Software. However, You may do so only on Your own behalf, and not on
behalf of any Contributor. You must make it absolutely clear that any
such warranty, support, indemnity, or liability obligation is offered by
You alone, and You hereby agree to indemnify every Contributor for any
liability incurred by such Contributor as a result of warranty, support,
indemnity or liability terms You offer. You may include additional
disclaimers of warranty and limitations of liability specific to any
jurisdiction.
4. Inability to Comply Due to Statute or Regulation
---------------------------------------------------
If it is impossible for You to comply with any of the terms of this
License with respect to some or all of the Covered Software due to
statute, judicial order, or regulation then You must: (a) comply with
the terms of this License to the maximum extent possible; and (b)
describe the limitations and the code they affect. Such description must
be placed in a text file included with all distributions of the Covered
Software under this License. Except to the extent prohibited by statute
or regulation, such description must be sufficiently detailed for a
recipient of ordinary skill to be able to understand it.
5. Termination
--------------
5.1. The rights granted under this License will terminate automatically
if You fail to comply with any of its terms. However, if You become
compliant, then the rights granted under this License from a particular
Contributor are reinstated (a) provisionally, unless and until such
Contributor explicitly and finally terminates Your grants, and (b) on an
ongoing basis, if such Contributor fails to notify You of the
non-compliance by some reasonable means prior to 60 days after You have
come back into compliance. Moreover, Your grants from a particular
Contributor are reinstated on an ongoing basis if such Contributor
notifies You of the non-compliance by some reasonable means, this is the
first time You have received notice of non-compliance with this License
from such Contributor, and You become compliant prior to 30 days after
Your receipt of the notice.
5.2. If You initiate litigation against any entity by asserting a patent
infringement claim (excluding declaratory judgment actions,
counter-claims, and cross-claims) alleging that a Contributor Version
directly or indirectly infringes any patent, then the rights granted to
You by any and all Contributors for the Covered Software under Section
2.1 of this License shall terminate.
5.3. In the event of termination under Sections 5.1 or 5.2 above, all
end user license agreements (excluding distributors and resellers) which
have been validly granted by You or Your distributors under this License
prior to termination shall survive termination.
************************************************************************
* *
* 6. Disclaimer of Warranty *
* ------------------------- *
* *
* Covered Software is provided under this License on an "as is" *
* basis, without warranty of any kind, either expressed, implied, or *
* statutory, including, without limitation, warranties that the *
* Covered Software is free of defects, merchantable, fit for a *
* particular purpose or non-infringing. The entire risk as to the *
* quality and performance of the Covered Software is with You. *
* Should any Covered Software prove defective in any respect, You *
* (not any Contributor) assume the cost of any necessary servicing, *
* repair, or correction. This disclaimer of warranty constitutes an *
* essential part of this License. No use of any Covered Software is *
* authorized under this License except under this disclaimer. *
* *
************************************************************************
************************************************************************
* *
* 7. Limitation of Liability *
* -------------------------- *
* *
* Under no circumstances and under no legal theory, whether tort *
* (including negligence), contract, or otherwise, shall any *
* Contributor, or anyone who distributes Covered Software as *
* permitted above, be liable to You for any direct, indirect, *
* special, incidental, or consequential damages of any character *
* including, without limitation, damages for lost profits, loss of *
* goodwill, work stoppage, computer failure or malfunction, or any *
* and all other commercial damages or losses, even if such party *
* shall have been informed of the possibility of such damages. This *
* limitation of liability shall not apply to liability for death or *
* personal injury resulting from such party's negligence to the *
* extent applicable law prohibits such limitation. Some *
* jurisdictions do not allow the exclusion or limitation of *
* incidental or consequential damages, so this exclusion and *
* limitation may not apply to You. *
* *
************************************************************************
8. Litigation
-------------
Any litigation relating to this License may be brought only in the
courts of a jurisdiction where the defendant maintains its principal
place of business and such litigation shall be governed by laws of that
jurisdiction, without reference to its conflict-of-law provisions.
Nothing in this Section shall prevent a party's ability to bring
cross-claims or counter-claims.
9. Miscellaneous
----------------
This License represents the complete agreement concerning the subject
matter hereof. If any provision of this License is held to be
unenforceable, such provision shall be reformed only to the extent
necessary to make it enforceable. Any law or regulation which provides
that the language of a contract shall be construed against the drafter
shall not be used to construe this License against a Contributor.
10. Versions of the License
---------------------------
10.1. New Versions
Mozilla Foundation is the license steward. Except as provided in Section
10.3, no one other than the license steward has the right to modify or
publish new versions of this License. Each version will be given a
distinguishing version number.
10.2. Effect of New Versions
You may distribute the Covered Software under the terms of the version
of the License under which You originally received the Covered Software,
or under the terms of any subsequent version published by the license
steward.
10.3. Modified Versions
If you create software not governed by this License, and you want to
create a new license for such software, you may create and use a
modified version of this License if you rename the license and remove
any references to the name of the license steward (except to note that
such modified license differs from this License).
10.4. Distributing Source Code Form that is Incompatible With Secondary
Licenses
If You choose to distribute Source Code Form that is Incompatible With
Secondary Licenses under the terms of this version of the License, the
notice described in Exhibit B of this License must be attached.
Exhibit A - Source Code Form License Notice
-------------------------------------------
This Source Code Form is subject to the terms of the Mozilla Public
License, v. 2.0. If a copy of the MPL was not distributed with this
file, You can obtain one at https://mozilla.org/MPL/2.0/.
If it is not possible or desirable to put the notice in a particular
file, then You may include the notice in a location (such as a LICENSE
file in a relevant directory) where a recipient would be likely to look
for such a notice.
You may add additional accurate notices of copyright ownership.
Exhibit B - "Incompatible With Secondary Licenses" Notice
---------------------------------------------------------
This Source Code Form is "Incompatible With Secondary Licenses", as
defined by the Mozilla Public License, v. 2.0.
+15
View File
@@ -0,0 +1,15 @@
services:
dtek_notif:
build:
context: .
dockerfile: Dockerfile
image: gcr.forust.xyz/forust/dtek-notif:latest
pull_policy: build
restart: unless-stopped
environment:
- TZ=Europe/Kyiv
dns:
- 1.1.1.1
- 8.8.8.8
networks:
- default
+748
View File
@@ -0,0 +1,748 @@
import asyncio
import contextlib
import logging
import os
from datetime import datetime, timedelta
import requests
from aiogram import Bot, Dispatcher
from aiogram.filters import Command
from aiogram.types import KeyboardButton, Message
from aiogram.utils.keyboard import ReplyKeyboardBuilder
from bs4 import BeautifulSoup
from dotenv import load_dotenv
# Загрузка переменных окружения
load_dotenv()
# Настройки
TELEGRAM_TOKEN = os.getenv('TELEGRAM_TOKEN', 'YOUR_TOKEN_HERE')
ALLOWED_CHAT_IDS = list(map(int, os.getenv('ALLOWED_CHAT_IDS', '').split(','))) if os.getenv('ALLOWED_CHAT_IDS') else []
CHECK_INTERVAL = int(os.getenv('CHECK_INTERVAL', '120'))
# Параметры для запроса
VOE_CITY_ID = int(os.getenv('VOE_CITY_ID', 'VOE_CITY_ID'))
VOE_STREET_ID = int(os.getenv('VOE_STREET_ID', 'VOE_STREET_ID'))
VOE_HOUSE_ID = int(os.getenv('VOE_HOUSE_ID', 'VOE_HOUSE_ID'))
# Настройка логирования
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
logger = logging.getLogger(__name__)
# Глобальные переменные
bot = Bot(token=TELEGRAM_TOKEN)
dp = Dispatcher()
last_schedule: list[dict] | None = None
last_notification_time: dict[str, datetime] = {}
# ============================================================================
# УТИЛИТЫ
# ============================================================================
def format_time_duration(minutes: int) -> str:
"""Форматирует время из минут в часы и минуты"""
hours = minutes // 60
mins = minutes % 60
if hours == 0:
return f'{mins}м'
elif mins == 0:
return f'{hours}ч'
return f'{hours}ч {mins}м'
def get_day_statistics(day_blocks: list[dict]) -> dict[str, int]:
"""Получает статистику по дню"""
total_minutes = 0
confirmed_minutes = 0
possible_minutes = 0
for block in day_blocks:
for half in [block['first_half'], block['second_half']]:
if half['status'] == 'off':
total_minutes += 30
if half['confirmed']:
confirmed_minutes += 30
else:
possible_minutes += 30
return {'total': total_minutes, 'confirmed': confirmed_minutes, 'possible': possible_minutes}
# ============================================================================
# ПАРСИНГ ДАННЫХ
# ============================================================================
def parse_html(html: str) -> list[dict]:
"""Парсит HTML с графиком отключений (логика от 15.11.2024)"""
soup = BeautifulSoup(html, 'html.parser')
cells = soup.select('.disconnection-detailed-table-cell.cell')
schedule = []
current_hour = 0
current_day = 0
for cell in cells:
if 'legend' in cell.get('class', []) or 'head' in cell.get('class', []):
continue
cell_classes = cell.get('class', [])
# ПРоверка статуса отключения на весь час
full_hour_off = 'has_disconnection' in cell_classes and 'full_hour' in cell_classes
hour_block = cell.select_one('.hour_block')
if not hour_block:
continue
# Проверка подтверждённости отключения для всего часа
cell_confirmed = None
if 'confirm_1' in cell_classes:
cell_confirmed = True
elif 'confirm_0' in cell_classes:
cell_confirmed = False
# Проверка половин часа
left = hour_block.select_one('.half.left')
right = hour_block.select_one('.half.right')
def parse_half(half, is_full_hour_off: bool, cell_confirmed: bool | None = None) -> dict:
"""Парсит половину часа"""
if not half:
return {'status': 'on', 'queue': None, 'confirmed': None}
half_classes = half.get('class', [])
# Если вся ячейка full_hour - используем статус ячейки
if is_full_hour_off:
return {'status': 'off', 'queue': None, 'confirmed': cell_confirmed}
# Определяем статус половины
if 'has_disconnection' in half_classes:
status = 'off'
elif 'no_disconnection' in half_classes:
status = 'on'
else:
status = 'on' # По умолчанию считаем включенным
# Если выключено - ищем подробности
queue = None
confirmed = None
if status == 'off':
disconnection_div = half.select_one('.disconnection')
if disconnection_div:
# Ищем номер черги в title
if disconnection_div.has_attr('title'):
title = disconnection_div['title']
if 'Номер черги' in title or 'Номер черги:' in title:
with contextlib.suppress(BaseException):
queue = title.split(':')[-1].strip()
# Определяем подтверждение
disc_classes = disconnection_div.get('class', [])
if 'disconnection_confirm_1' in disc_classes:
confirmed = True
elif 'disconnection_confirm_0' in disc_classes:
confirmed = False
return {'status': status, 'queue': queue, 'confirmed': confirmed}
first_half_data = parse_half(left, full_hour_off, cell_confirmed)
second_half_data = parse_half(right, full_hour_off, cell_confirmed)
schedule.append(
{
'hour': current_hour,
'day': current_day,
'first_half': first_half_data,
'second_half': second_half_data,
}
)
current_hour += 1
if current_hour >= 24:
current_hour = 0
current_day += 1
return schedule
def get_voe_html(city_id: int, street_id: int, house_id: int) -> str:
"""Получает HTML с сайта VOE"""
url = 'https://www.voe.com.ua/disconnection/detailed?ajax_form=1&_wrapper_format=drupal_ajax'
headers = {
'Content-Type': 'application/x-www-form-urlencoded; charset=UTF-8',
'X-Requested-With': 'XMLHttpRequest',
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36',
}
data = {
'search_type': 0,
'city_id': city_id,
'street_id': street_id,
'house_id': house_id,
'form_build_id': 'form-Irv5aHw1R2FT_Ik2apyHOZ47hTH5xPNH_LQnBrmpSTc',
'form_id': 'disconnection_detailed_search_form',
'_triggering_element_name': 'search',
'_triggering_element_value': 'Показати',
'_drupal_ajax': 1,
}
try:
response = requests.post(url, headers=headers, data=data, timeout=10)
response.raise_for_status()
resp_json = response.json()
insert_html = next((item['data'] for item in resp_json if item.get('command') == 'insert'), None)
if not insert_html:
raise ValueError('HTML не найден в ответе')
return insert_html
except requests.exceptions.RequestException as e:
logger.error(f'Ошибка запроса VOE: {e}')
raise
# ============================================================================
# ФОРМАТИРОВАНИЕ СООБЩЕНИЙ
# ============================================================================
def get_main_keyboard():
"""Создает главную клавиатуру"""
builder = ReplyKeyboardBuilder()
builder.row(KeyboardButton(text='📊 Графік'), KeyboardButton(text='🔄 Оновити'))
builder.row(KeyboardButton(text='📅 Сьогодні'), KeyboardButton(text='📅 Завтра'))
builder.row(KeyboardButton(text='ℹ️ Про бота'))
return builder.as_markup(resize_keyboard=True)
def format_schedule_message(schedule: list[dict], days_to_show: int = 2) -> str:
"""Форматирует полный график на несколько дней"""
lines = [
'⚡️ <b>Графік відключень світла</b>',
f'🕐 Оновлено: {datetime.now().strftime("%d.%m.%Y %H:%M:%S")}',
'─' * 30,
'',
]
start_date = datetime.now()
for day in range(min(days_to_show, 2)):
day_blocks = [b for b in schedule if b['day'] == day]
if not day_blocks:
continue
date_str = (start_date + timedelta(days=day)).strftime('%d.%m.%Y')
day_name = '🌅 <b>Сьогодні</b>' if day == 0 else '🌄 <b>Завтра</b>'
lines.append(f'{day_name} ({date_str})')
# Статистика
stats = get_day_statistics(day_blocks)
if stats['total'] > 0:
lines.append(f'⏱ Всього: <code>{format_time_duration(stats["total"])}</code>')
if stats['confirmed'] > 0:
lines.append(f'🔴 Підтверджено: <code>{format_time_duration(stats["confirmed"])}</code>')
if stats['possible'] > 0:
lines.append(f'🟠 Можливо: <code>{format_time_duration(stats["possible"])}</code>')
else:
lines.append('🟢 <b>Відключень немає!</b>')
lines.append('')
# Детальный список отключений
disconnections = []
current_status = None
start_time = None
current_confirmed = None
current_queue = None
for block in day_blocks:
hour = block['hour']
for half_idx, half in enumerate([block['first_half'], block['second_half']]):
time_str = f'{hour:02d}:00' if half_idx == 0 else f'{hour:02d}:30'
if half['status'] == 'off':
if current_status != 'off':
start_time = time_str
current_confirmed = half['confirmed']
current_queue = half['queue']
current_status = 'off'
else:
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч{current_queue})' if current_queue else ''
disconnections.append(f'{icon} <code>{start_time} - {time_str}</code>{queue_text}')
current_status = half['status']
# Если день закончился на отключении
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч{current_queue})' if current_queue else ''
next_hour = (day_blocks[-1]['hour'] + 1) % 24
end_time = f'{next_hour:02d}:00'
disconnections.append(f'{icon} <code>{start_time} - {end_time}</code>{queue_text}')
if disconnections:
for idx, disc in enumerate(disconnections, 1):
lines.append(f'{idx}. {disc}')
lines.append('')
lines.append('<i>🔴 = підтверджено • 🟠 = можливо • 🟢 = світло</i>')
return '\n'.join(lines)
def format_single_day_schedule(schedule: list[dict], day: int) -> str:
"""Форматирует график на один день"""
day_blocks = [b for b in schedule if b['day'] == day]
if not day_blocks:
return '❌ Немає даних для цього дня'
start_date = datetime.now()
date_str = (start_date + timedelta(days=day)).strftime('%d.%m.%Y')
day_name = '🟠 <b>Сьогодні</b>' if day == 0 else '🔶 <b>Завтра</b>'
lines = [f'{day_name} • {date_str}', '']
# Статистика
lines.append('<b>📊 Статистика</b>')
stats = get_day_statistics(day_blocks)
if stats['total'] == 0:
lines.append('└ 🟢 <b>Відключень немає!</b>')
else:
total_time = format_time_duration(stats['total'])
lines.append(f'├ ⏱ Всього: <code>{total_time}</code>')
if stats['confirmed'] > 0:
confirmed_time = format_time_duration(stats['confirmed'])
lines.append(f'├ 🔴 Підтверджено: <code>{confirmed_time}</code>')
if stats['possible'] > 0:
possible_time = format_time_duration(stats['possible'])
lines.append(f'└ 🟠 Можливо: <code>{possible_time}</code>')
else:
lines.append('└ 🟢 Решта часу світло')
lines.append('')
# Детальный список отключений
disconnections = []
current_status = None
start_time = None
current_confirmed = None
current_queue = None
for block in day_blocks:
hour = block['hour']
for half_idx, half in enumerate([block['first_half'], block['second_half']]):
time_str = f'{hour:02d}:00' if half_idx == 0 else f'{hour:02d}:30'
if half['status'] == 'off':
if current_status != 'off':
start_time = time_str
current_confirmed = half['confirmed']
current_queue = half['queue']
current_status = 'off'
else:
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч.{current_queue})' if current_queue else ''
disconnections.append(f'{icon} <code>{start_time} - {time_str}</code>{queue_text}')
current_status = half['status']
# Если день закончился на отключении
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч.{current_queue})' if current_queue else ''
next_hour = (day_blocks[-1]['hour'] + 1) % 24
end_time = f'{next_hour:02d}:00'
disconnections.append(f'{icon} <code>{start_time} - {end_time}</code>{queue_text}')
if disconnections:
lines.append('<b>⚡️ Розклад відключень</b>')
for idx, disc in enumerate(disconnections, 1):
lines.append(f'{idx}. {disc}')
lines.append('')
lines.append('<i>🔴 підтверджено • 🟠 можливо • 🟢 світло</i>')
return '\n'.join(lines)
def schedules_differ(old_schedule: list[dict] | None, new_schedule: list[dict] | None) -> bool:
"""Проверяет отличия между графиками"""
if old_schedule is None or new_schedule is None:
return True
if len(old_schedule) != len(new_schedule):
return True
for old, new in zip(old_schedule, new_schedule, strict=False):
if old['day'] >= 2:
break
if old['first_half'] != new['first_half'] or old['second_half'] != new['second_half']:
return True
return False
# ============================================================================
# УВЕДОМЛЕНИЯ
# ============================================================================
async def send_to_all_users(message_text: str, parse_mode: str = 'HTML'):
"""Отправляет сообщение всем пользователям"""
if not ALLOWED_CHAT_IDS:
logger.warning('Нет допущенных ID чатов для отправки уведомлений')
return
for chat_id in ALLOWED_CHAT_IDS:
try:
await bot.send_message(chat_id, message_text, parse_mode=parse_mode)
logger.info(f'✅ Сообщение отправлено пользователю {chat_id}')
except Exception as e:
logger.error(f'❌ Ошибка отправки пользователю {chat_id}: {e}')
await asyncio.sleep(0.5)
async def check_schedule():
"""Проверяет график и отправляет уведомления"""
global last_schedule
try:
logger.info('🔍 Проверка графика...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
new_schedule = parse_html(html)
if schedules_differ(last_schedule, new_schedule):
logger.info('✨ Обнаружены изменения!')
message = format_schedule_message(new_schedule, days_to_show=2)
if last_schedule is not None:
await send_to_all_users(f'🔄 <b>Графік оновлено!</b>\n\n{message}')
last_schedule = new_schedule
else:
logger.info('✓ Графік без змін')
except Exception as e:
logger.error(f'❌ Ошибка при проверке графика: {e}')
async def check_upcoming_disconnections():
"""Проверяет предстоящие события и отправляет предупреждения за 5 минут"""
global last_notification_time
if last_schedule is None:
return
now = datetime.now()
today_blocks = [b for b in last_schedule if b['day'] == 0]
# Создаем список всех переходов (off -> on или on -> off)
transitions = []
prev_status = None
for block in today_blocks:
hour = block['hour']
for half_idx, half in enumerate([block['first_half'], block['second_half']]):
minute = 0 if half_idx == 0 else 30
time_str = f'{hour:02d}:{minute:02d}'
current_status = half['status']
# Если статус изменился - это переход
if prev_status is not None and prev_status != current_status:
transitions.append(
{
'hour': hour,
'minute': minute,
'time_str': time_str,
'from_status': prev_status,
'to_status': current_status,
'confirmed': half.get('confirmed'),
'queue': half.get('queue'),
}
)
prev_status = current_status
# Проверяем переходы
for transition in transitions:
event_time = now.replace(hour=transition['hour'], minute=transition['minute'], second=0, microsecond=0)
time_until = (event_time - now).total_seconds() / 60
notification_key = f'{transition["hour"]}:{transition["minute"]}_{transition["to_status"]}'
# Если за 5 минут до события (±1 минута) и еще не отправляли
if 4 <= time_until <= 6:
# Проверяем, не отправляли ли уже уведомление сегодня
if notification_key in last_notification_time:
last_notif_time = last_notification_time[notification_key]
if last_notif_time.date() == now.date():
continue # Уже отправляли сегодня
# Переход на ОТКЛЮЧЕНИЕ (on -> off)
if transition['from_status'] == 'on' and transition['to_status'] == 'off':
icon = '🔴' if transition['confirmed'] else '🟠'
status = 'підтверджено' if transition['confirmed'] else 'можливе'
queue_info = f' (Черга {transition["queue"]})' if transition['queue'] else ''
warning = (
f'⚠️ <b>УВАГА! ВІДКЛЮЧЕННЯ</b>\n\n'
f'Через ~5 хвилин\n'
f'Час: <code>{transition["time_str"]}</code>\n'
f'Статус: {icon} {status}{queue_info}'
)
await send_to_all_users(warning)
last_notification_time[notification_key] = now
logger.info(f'📢 Відправлено попередження про ВІДКЛЮЧЕННЯ в {transition["time_str"]}')
# Переход на ВКЛЮЧЕНИЕ (off -> on)
elif transition['from_status'] == 'off' and transition['to_status'] == 'on':
warning = (
f'✅ <b>УВАГА! ВКЛЮЧЕННЯ</b>\n\n'
f'Через ~5 хвилин буде світло\n'
f'Час: <code>{transition["time_str"]}</code>'
)
await send_to_all_users(warning)
last_notification_time[notification_key] = now
logger.info(f'📢 Відправлено попередження про ВКЛЮЧЕННЯ в {transition["time_str"]}')
async def monitoring_loop():
"""Основной цикл мониторинга"""
await check_schedule()
while True:
try:
await asyncio.sleep(CHECK_INTERVAL)
await check_schedule()
await check_upcoming_disconnections()
except Exception as e:
logger.error(f'Ошибка в цикле мониторинга: {e}')
await asyncio.sleep(5)
# ============================================================================
# ОБРАБОТЧИКИ КОМАНД
# ============================================================================
@dp.message(Command('start'))
async def cmd_start(message: Message):
"""Обработчик /start"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу до цього бота.')
return
await message.answer(
'👋 <b>Ласкаво просимо!</b>\n\n'
'🤖 <b>Бот для моніторингу графіку відключень світла</b>\n\n'
'✨ <b>Можливості:</b>\n'
'• 📊 Перегляд графіку на сьогодні і завтра\n'
'• 🔔 Автоматичні сповіщення за 5 хвилин до подій\n'
'• 🔄 Моніторинг змін графіку\n\n'
'Використовуйте кнопки нижче 👇',
parse_mode='HTML',
reply_markup=get_main_keyboard(),
)
@dp.message(lambda msg: msg.text == 'ℹ️ Про бота')
async def cmd_info(message: Message):
"""Показывает информацию о боте"""
if message.chat.id not in ALLOWED_CHAT_IDS:
return
await message.answer(
'<b>ℹ️ Про бота</b>\n\n'
'🚀 <b>Версія:</b> 2.2 (Стабільна)\n\n'
'📝 <b>Реліз-ноути:</b>\n'
'├ 15.11.2024: Адаптація під оновлену логіку сайту VOE\n'
'├ Виправлено парсинг half.left та half.right\n'
'├ Покращено визначення підтвердження відключень\n'
'└ Оптимізовано обробку статусу для всієї години\n\n'
'⚡ <b>Функціональність:</b>\n'
'├ Моніторинг графіку 24/7\n'
'├ Сповіщення за 5 хвилин\n'
'├ Детальна статистика дня\n'
'└ Красива візуалізація\n\n'
'🔐 <b>Безпека:</b> Використовуються .env файли\n'
'💾 <b>Джерело:</b> voe.com.ua',
parse_mode='HTML',
reply_markup=get_main_keyboard(),
)
@dp.message(Command('schedule'))
async def cmd_schedule(message: Message):
"""Показывает полный график"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('⏳ Завантаження графіку...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
schedule = parse_html(html)
text = format_schedule_message(schedule, days_to_show=2)
await message.answer(text, parse_mode='HTML', reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message(Command('today'))
async def cmd_today(message: Message):
"""Показывает график на сегодня"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('⏳ Завантаження графіку сьогодні...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
schedule = parse_html(html)
text = format_single_day_schedule(schedule, 0)
await message.answer(text, parse_mode='HTML', reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message(Command('tomorrow'))
async def cmd_tomorrow(message: Message):
"""Показывает график на завтра"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('⏳ Завантаження графіку завтра...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
schedule = parse_html(html)
text = format_single_day_schedule(schedule, 1)
await message.answer(text, parse_mode='HTML', reply_markup=get_main_keyboard())
# await message.answer("❌ Функція тимчасово недоступна. Чекаємо на оновлення сайту", parse_mode="HTML", reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message(Command('check'))
async def cmd_check(message: Message):
"""Принудительная проверка графика"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('🔄 <b>Перевіряю графік...</b>', parse_mode='HTML')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
new_schedule = parse_html(html)
prefix = (
'✅ <b>Знайдено зміни!</b>\n\n'
if schedules_differ(last_schedule, new_schedule)
else '✓ <b>Графік без змін</b>\n\n'
)
result = prefix + format_schedule_message(new_schedule, days_to_show=2)
await message.answer(result, parse_mode='HTML', reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message()
async def handle_text(message: Message):
"""Обработчик текстовых сообщений и кнопок"""
if message.chat.id not in ALLOWED_CHAT_IDS:
return
text = message.text
# Кнопка "Графік"
if text == '📊 Графік':
await cmd_schedule(message)
# Кнопка "Сьогодні"
elif text == '📅 Сьогодні':
await cmd_today(message)
# Кнопка "Завтра"
elif text == '📅 Завтра':
await cmd_tomorrow(message)
# Кнопка "Оновити"
elif text == '🔄 Оновити':
await cmd_check(message)
# Кнопка "Про бота"
elif text == 'ℹ️ Про бота':
await cmd_info(message)
# Неизвестная команда
else:
await message.answer(
'❓ <b>Команда не розпізнана</b>\n\n'
'Використовуйте кнопки на клавіатурі або команди:\n'
'/start • /today • /tomorrow • /schedule • /check',
parse_mode='HTML',
reply_markup=get_main_keyboard(),
)
# ============================================================================
# ГЛАВНАЯ ФУНКЦИЯ
# ============================================================================
async def main():
"""Главная функция"""
logger.info('=' * 50)
logger.info('ЗАПУСК БОТА V2.2 (stable 2.2, 15.11.2025)')
logger.info('=' * 50)
if not TELEGRAM_TOKEN or os.getenv('TELEGRAM_TOKEN', 'YOUR_TOKEN_HERE') == TELEGRAM_TOKEN:
logger.error('❌ TELEGRAM_TOKEN не конфігурований! Напишіть токен в .env файл')
return
if not ALLOWED_CHAT_IDS:
logger.error('❌ ALLOWED_CHAT_IDS не конфігуровані! Напишіть ID в .env файл')
return
logger.info(f'📌 Allowed chat ids: {ALLOWED_CHAT_IDS}')
logger.info(f'⏱ Інтервал перевірки: {CHECK_INTERVAL} сек')
logger.info('=' * 50)
# Запускаем мониторинг
monitoring_task = asyncio.create_task(monitoring_loop())
try:
await dp.start_polling(bot)
except KeyboardInterrupt:
logger.info('⏹ Бот зупинений користувачем')
finally:
monitoring_task.cancel()
await bot.session.close()
logger.info('✓ Підключення закрито')
if __name__ == '__main__':
try:
asyncio.run(main())
except KeyboardInterrupt:
logger.info('⏹ Завершено')
+7
View File
@@ -0,0 +1,7 @@
[project]
name = "dtek-notif"
version = "0.1.0"
description = "Add your description here"
readme = "README.md"
requires-python = ">=3.13"
dependencies = []
+5
View File
@@ -0,0 +1,5 @@
requests>=2.31.0
beautifulsoup4>=4.12.0
aiogram>=3.3.0
python-dotenv>=1.0.0
aiohttp>=3.9.0
-1
View File
@@ -1 +0,0 @@
1.56.0
-52
View File
@@ -1,52 +0,0 @@
# EDU session keeper and Telegram bot
Keeps an EDU login session in Redis and sends Telegram notifications for new webinars. The bot also serves diary and schedule commands.
`phpsessid-bot/` logs into EDU and publishes `EDU_PHPSESSID` in Redis.
`webinar-checker/` uses that cookie through a remote Playwright browser and stores
subscribers, language preferences, and webinar history in Redis.
Kubernetes runs in `edu-master`, with Redis data in `redis-data-pvc`.
`service.yaml`, `servicemonitor.yaml`, and `alerts.yaml` expose and monitor the
checker's metrics on port 8000. Its `/health` endpoint reflects recent checks.
## Configuration
Use the keys in `k8s/secrets.yaml.example` as the reference. The committed Compose
`.env.example` has stale names until `fix/session-keeper-reliability` is merged.
The code reads:
| Variable | Purpose |
| ----------------------------------------------------- | ------------------------------------------------ |
| `KEEPER_LOGIN`, `KEEPER_PASSWORD` | EDU login credentials. |
| `KEEPER_INTERVAL` | Session refresh interval in minutes; default 10. |
| `EDU_URL_BASE` | EDU site origin. |
| `EDU_URL_LOGIN`, `EDU_URL_COURSES`, `EDU_URL_WEBINAR` | Paths under that origin. |
| `WEBINAR_TELEGRAM_TOKEN`, `WEBINAR_ADMIN_ID` | Telegram bot and administrator. |
| `WEBINAR_CHECK_INTERVAL` | Checker interval in seconds; default 60. |
| `REDIS_HOST`, `REDIS_PORT` | Redis connection. |
| `PLAYWRIGHT_WS` | Remote browser WebSocket endpoint. |
Set the keeper keys explicitly in the Compose `.env`. Keep the Playwright Python
package, browser image, server command, and `PLAYWRIGHT_VERSION` file on matching
versions. The two Python images are built and published by CI.
## Bot use
Start a private chat with `/start` to subscribe. `/stop`, `/language`, `/diary`,
`/schedule`, and `/setclass` manage subscriptions and school views. The
administrator can manage the whitelist with `/adduser` and `/removeuser`.
Back up Redis if subscriber settings and notification history matter. Session
cookies and Telegram tokens are credentials; keep them out of shared logs.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n edu-master
kubectl get events -n edu-master --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+4 -4
View File
@@ -1,6 +1,6 @@
services:
redis:
image: redis:8.10.2-alpine
image: redis:8.10.1-alpine
restart: unless-stopped
volumes:
- redis-data:/data
@@ -11,13 +11,13 @@ services:
retries: 5
playwright-service:
image: mcr.microsoft.com/playwright:v1.56.0-jammy
image: mcr.microsoft.com/playwright:v1.63.0-jammy
restart: unless-stopped
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
session-keeper:
build: ./phpsessid-bot
image: gcr.forust.xyz/forust/session-keeper:prod
image: gcr.forust.xyz/forust/session-keeper:latest
pull_policy: build
env_file: .env
restart: unless-stopped
@@ -33,7 +33,7 @@ services:
webinar-checker:
build: ./webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:prod
image: gcr.forust.xyz/forust/webinar-checker:latest
pull_policy: build
env_file: .env
restart: unless-stopped
-96
View File
@@ -1,96 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: edu-master-webinar
namespace: edu-master
labels:
release: prometheus-stack
spec:
groups:
- name: edu_master.webinar
rules:
# No successful webinar check for 5m (~2-3 missed 2-min checks).
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
# The last_success > 0 guard is mandatory: checker.py initialises
# last_success to 0, so without it `time() - 0` equals the current epoch
# and humanizeDuration renders ~20722d on every pod restart. Keep the
# duration expression on the left so $value stays the real gap.
- alert: WebinarCheckerNoSuccessfulCheck
expr: |
((time() - webinar_check_last_success_timestamp_seconds) > 300)
and (webinar_check_last_success_timestamp_seconds > 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker has no successful check for 5m"
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Checks are running but none has ever succeeded since pod start.
# Split out from the rule above so a zeroed gauge never feeds
# humanizeDuration.
- alert: WebinarCheckerNeverSucceeded
expr: |
(webinar_check_last_success_timestamp_seconds == 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker has never completed a successful check"
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
- alert: WebinarCheckerConsecutiveFailures
expr: |
webinar_check_consecutive_failures >= 3
for: 5m
labels:
severity: critical
annotations:
summary: "Webinar checker failing consecutively"
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / playwright error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
- alert: WebinarCheckerScrapeDown
expr: |
absent(webinar_check_last_run_timestamp_seconds) == 1
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker metrics missing"
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
- alert: EduPhpsessidMissing
expr: |
edu_phpsessid_present == 0
for: 10m
labels:
severity: critical
annotations:
summary: "EDU_PHPSESSID missing"
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker and playwright deployments unavailable.
- alert: WebinarCheckerDeploymentDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker deployment unavailable"
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
- alert: PlaywrightServiceDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="playwright-service", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Playwright service unavailable"
description: "edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
+1 -13
View File
@@ -10,8 +10,6 @@ spec:
selector:
matchLabels:
app: edu-master-playwright
strategy:
type: Recreate
template:
metadata:
labels:
@@ -19,18 +17,8 @@ spec:
spec:
containers:
- name: playwright
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
image: mcr.microsoft.com/playwright:v1.56.0-jammy
image: mcr.microsoft.com/playwright:v1.63.0-jammy
imagePullPolicy: IfNotPresent
# p95 412M, max 478M over 7 days, no limit before. Request is set at p95
# so the pod is not an eviction candidate; the limit stays above 2x the
# request because browser page lifetimes are unpredictable.
resources:
requests:
cpu: "200m"
memory: "416Mi"
limits:
memory: "1Gi"
command:
- npx
- -y
+3 -3
View File
@@ -18,7 +18,7 @@ spec:
spec:
containers:
- name: redis
image: redis:8.10.2-alpine
image: redis:8.10.1-alpine
imagePullPolicy: IfNotPresent
ports:
- containerPort: 6379
@@ -28,10 +28,10 @@ spec:
resources:
requests:
cpu: 25m
memory: 32Mi
memory: 64Mi
limits:
cpu: 250m
memory: 128Mi
memory: 256Mi
readinessProbe:
exec:
command: ["redis-cli", "ping"]
-2
View File
@@ -21,8 +21,6 @@ stringData:
WEBINAR_TELEGRAM_TOKEN: ""
WEBINAR_ADMIN_ID: ""
WEBINAR_CHECK_INTERVAL: "60"
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
METRICS_PORT: "8000"
# Database
REDIS_HOST: "redis"
REDIS_PORT: "6379"
-15
View File
@@ -1,15 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
selector:
app: edu-master-webinar-checker
ports:
- name: metrics
port: 8000
targetPort: metrics
protocol: TCP
-16
View File
@@ -1,16 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: webinar-checker
namespace: edu-master
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: edu-master-webinar-checker
endpoints:
- port: metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
+5 -6
View File
@@ -10,8 +10,6 @@ spec:
selector:
matchLabels:
app: edu-master-session-keeper
strategy:
type: Recreate
template:
metadata:
labels:
@@ -19,7 +17,7 @@ spec:
spec:
initContainers:
- name: wait-redis
image: redis:8.10.2-alpine
image: redis:8.10.1-alpine
command:
- /bin/sh
- -ec
@@ -33,17 +31,18 @@ spec:
echo "redis is ready"
containers:
- name: session-keeper
image: gcr.forust.xyz/forust/session-keeper:prod
image: gcr.forust.xyz/forust/session-keeper:latest
imagePullPolicy: Always
envFrom:
- secretRef:
name: edu-master-secrets
resources:
requests:
cpu: 25m
memory: 32Mi
memory: 96Mi
limits:
cpu: 250m
memory: 128Mi
memory: 256Mi
readinessProbe:
exec:
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
+5 -18
View File
@@ -10,8 +10,6 @@ spec:
selector:
matchLabels:
app: edu-master-webinar-checker
strategy:
type: Recreate
template:
metadata:
labels:
@@ -21,7 +19,7 @@ spec:
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started
initContainers:
- name: wait-deps
image: redis:8.10.2-alpine
image: redis:8.10.1-alpine
command:
- /bin/sh
- -ec
@@ -47,19 +45,8 @@ spec:
echo "playwright ok"
containers:
- name: webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:prod
ports:
- name: metrics
containerPort: 8000
protocol: TCP
readinessProbe:
httpGet:
path: /health
port: metrics
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 12
initialDelaySeconds: 10
image: gcr.forust.xyz/forust/webinar-checker:latest
imagePullPolicy: Always
envFrom:
- secretRef:
name: edu-master-secrets
@@ -69,7 +56,7 @@ spec:
resources:
requests:
cpu: "50m"
memory: "192Mi"
memory: "128Mi"
limits:
cpu: "600m"
memory: "384Mi"
memory: "512Mi"
+2 -5
View File
@@ -2,11 +2,8 @@ FROM python:3.11-slim
WORKDIR /app
# renovate: datasource=pypi depName=playwright versioning=pep440
ARG PLAYWRIGHT_VERSION=1.56.0
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
# Install dependencies
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==1.56.0 redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
COPY checker.py .
+144 -277
View File
@@ -1,15 +1,12 @@
import asyncio
import contextlib
import json
import logging
import os
import re
import tempfile
import threading
import time
from datetime import datetime, timedelta
from html import escape
from http.server import BaseHTTPRequestHandler, HTTPServer
import redis
from playwright.async_api import async_playwright
@@ -51,106 +48,6 @@ USER_AGENT = _env(
)
WEBINAR_TELEGRAM_TOKEN = _env('WEBINAR_TELEGRAM_TOKEN')
ADMIN_ID = int(_env('WEBINAR_ADMIN_ID', '0'))
METRICS_PORT = int(_env('METRICS_PORT', '8000'))
# --- Prometheus metrics (stdlib only, no extra deps) ---
# Scraped by prometheus-stack via ServiceMonitor (edu_master/k8s/servicemonitor.yaml).
# Critical alerts in edu_master/k8s/alerts.yaml fire to Telegram via Alertmanager.
_METRICS_LOCK = threading.Lock()
_METRICS = {
'last_run': 0.0, # Unix ts of last check start
'last_success': 0.0, # Unix ts of last successful check
'last_duration': 0.0, # Duration of last check in seconds
'success_total': 0,
'failure_total': 0,
'consecutive_failures': 0,
'phpsessid_present': 1, # 1 if EDU_PHPSESSID found in redis, else 0
}
def _metric_check_start():
with _METRICS_LOCK:
_METRICS['last_run'] = time.time()
def _metric_check_ok(duration: float):
now = time.time()
with _METRICS_LOCK:
_METRICS['last_success'] = now
_METRICS['last_duration'] = duration
_METRICS['success_total'] += 1
_METRICS['consecutive_failures'] = 0
_METRICS['phpsessid_present'] = 1
def _metric_check_fail(duration: float, phpsessid_missing: bool = False):
with _METRICS_LOCK:
_METRICS['last_duration'] = duration
_METRICS['failure_total'] += 1
_METRICS['consecutive_failures'] += 1
_METRICS['phpsessid_present'] = 0 if phpsessid_missing else 1
def _metrics_render() -> bytes:
with _METRICS_LOCK:
m = dict(_METRICS)
lines = [
'# HELP webinar_check_last_run_timestamp_seconds Unix timestamp of last webinar check start.',
'# TYPE webinar_check_last_run_timestamp_seconds gauge',
f'webinar_check_last_run_timestamp_seconds {m["last_run"]}',
'# HELP webinar_check_last_success_timestamp_seconds Unix timestamp of last successful webinar check.',
'# TYPE webinar_check_last_success_timestamp_seconds gauge',
f'webinar_check_last_success_timestamp_seconds {m["last_success"]}',
'# HELP webinar_check_last_duration_seconds Duration of last webinar check in seconds.',
'# TYPE webinar_check_last_duration_seconds gauge',
f'webinar_check_last_duration_seconds {m["last_duration"]}',
'# HELP webinar_check_success_total Total successful webinar checks.',
'# TYPE webinar_check_success_total counter',
f'webinar_check_success_total {m["success_total"]}',
'# HELP webinar_check_failure_total Total failed webinar checks (timeout, playwright error, page error).',
'# TYPE webinar_check_failure_total counter',
f'webinar_check_failure_total {m["failure_total"]}',
'# HELP webinar_check_consecutive_failures Consecutive failed webinar checks (reset on success).',
'# TYPE webinar_check_consecutive_failures gauge',
f'webinar_check_consecutive_failures {m["consecutive_failures"]}',
'# HELP edu_phpsessid_present 1 if EDU_PHPSESSID exists in redis, 0 otherwise.',
'# TYPE edu_phpsessid_present gauge',
f'edu_phpsessid_present {m["phpsessid_present"]}',
]
return ('\n'.join(lines) + '\n').encode()
class _MetricsHandler(BaseHTTPRequestHandler):
def do_GET(self):
if self.path == '/metrics':
body = _metrics_render()
self.send_response(200)
self.send_header('Content-Type', 'text/plain; version=0.0.4')
self.send_header('Content-Length', str(len(body)))
self.end_headers()
self.wfile.write(body)
elif self.path in ('/healthz', '/health'):
body = b'ok\n'
self.send_response(200)
self.send_header('Content-Type', 'text/plain')
self.send_header('Content-Length', str(len(body)))
self.end_headers()
self.wfile.write(body)
else:
self.send_response(404)
self.end_headers()
def log_message(self, *args):
pass # keep bot logs clean
def start_metrics_server(port: int = METRICS_PORT):
server = HTTPServer(('0.0.0.0', port), _MetricsHandler) # noqa: S104 - k8s ServiceMonitor scrapes pod IP
thread = threading.Thread(target=server.serve_forever, name='metrics-server', daemon=True)
thread.start()
logger.info(f'Metrics server listening on :{port}/metrics')
return server
# Redis Keys
KEY_WHITELIST = 'bot:whitelist'
@@ -700,66 +597,59 @@ async def _collect_event_times(page) -> dict:
async def fetch_diary_data(phpsessid: str) -> dict | None:
logger.info('Fetching diary data via Playwright...')
try:
async with asyncio.timeout(60):
async with async_playwright() as p:
browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15)
async with async_playwright() as p:
browser = await p.chromium.connect(PLAYWRIGHT_WS)
try:
context_browser = await browser.new_context(user_agent=USER_AGENT)
await context_browser.add_cookies(
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}]
)
page = await context_browser.new_page()
try:
context_browser = await browser.new_context(user_agent=USER_AGENT)
await context_browser.add_cookies(
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}]
)
page = await context_browser.new_page()
await page.goto(DIARY_URL, wait_until='domcontentloaded')
await page.wait_for_selector('table.calendar', timeout=10000)
await page.wait_for_timeout(1500)
try:
await asyncio.wait_for(page.goto(DIARY_URL, wait_until='domcontentloaded'), timeout=30)
await page.wait_for_selector('table.calendar', timeout=10000)
await page.wait_for_timeout(1500)
table_html = await page.evaluate("""
() => {
const t = document.querySelector('table.calendar');
return t ? t.outerHTML : null;
}
""")
if not table_html:
logger.error('table.calendar not found in DOM')
return None
# Debug: save HTML for troubleshooting
with contextlib.suppress(Exception), open('/tmp/diary_debug.html', 'w', encoding='utf-8') as f: # noqa: S108
f.write(table_html)
month_text, days = _parse_calendar_html(table_html)
# Read event times by opening each event's AJAX popup.
times_by_id = await _collect_event_times(page)
if times_by_id:
for day_data in days.values():
for ev in day_data.get('events', []):
eid = ev.get('id')
if eid and eid in times_by_id:
ev['time'] = times_by_id[eid]
logger.info(
f'Diary parsed: month={month_text!r}, days_with_events={sum(1 for d in days.values() if d["events"])}/{len(days)}'
)
return {'monthFullText': month_text, 'days': days}
except Exception as e:
logger.error(f'Error parsing diary: {e}')
table_html = await page.evaluate("""
() => {
const t = document.querySelector('table.calendar');
return t ? t.outerHTML : null;
}
""")
if not table_html:
logger.error('table.calendar not found in DOM')
return None
finally:
with contextlib.suppress(Exception):
await asyncio.wait_for(page.close(), timeout=5)
with contextlib.suppress(Exception):
await asyncio.wait_for(context_browser.close(), timeout=5)
# Debug: save HTML for troubleshooting
with contextlib.suppress(Exception), open('/tmp/diary_debug.html', 'w', encoding='utf-8') as f: # noqa: S108
f.write(table_html)
month_text, days = _parse_calendar_html(table_html)
# Read event times by opening each event's AJAX popup.
times_by_id = await _collect_event_times(page)
if times_by_id:
for day_data in days.values():
for ev in day_data.get('events', []):
eid = ev.get('id')
if eid and eid in times_by_id:
ev['time'] = times_by_id[eid]
logger.info(
f'Diary parsed: month={month_text!r}, days_with_events={sum(1 for d in days.values() if d["events"])}/{len(days)}'
)
return {'monthFullText': month_text, 'days': days}
except Exception as e:
logger.error(f'Error parsing diary: {e}')
return None
finally:
with contextlib.suppress(Exception):
await asyncio.wait_for(browser.close(), timeout=5)
except TimeoutError:
logger.error('Diary fetch timed out (60s)')
return None
await page.close()
await context_browser.close()
finally:
await browser.close()
except Exception as e:
logger.error(f'Playwright error in diary fetch: {e}')
return None
@@ -1041,55 +931,48 @@ def _parse_schedule_html(table_html: str) -> dict:
async def fetch_schedule_data(phpsessid: str) -> dict | None:
logger.info('Fetching schedule data via Playwright...')
try:
async with asyncio.timeout(60):
async with async_playwright() as p:
browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15)
async with async_playwright() as p:
browser = await p.chromium.connect(PLAYWRIGHT_WS)
try:
context_browser = await browser.new_context(user_agent=USER_AGENT)
await context_browser.add_cookies(
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}]
)
page = await context_browser.new_page()
try:
context_browser = await browser.new_context(user_agent=USER_AGENT)
await context_browser.add_cookies(
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}]
)
page = await context_browser.new_page()
await page.goto(SCHEDULE_URL, wait_until='domcontentloaded')
await page.wait_for_selector('table.schedule-table', timeout=10000)
await page.wait_for_timeout(1500)
try:
await asyncio.wait_for(page.goto(SCHEDULE_URL, wait_until='domcontentloaded'), timeout=30)
await page.wait_for_selector('table.schedule-table', timeout=10000)
await page.wait_for_timeout(1500)
table_html = await page.evaluate("""
() => {
const t = document.querySelector('table.schedule-table');
return t ? t.outerHTML : null;
}
""")
if not table_html:
logger.error('table.schedule-table not found in DOM')
return None
debug_path = os.path.join(tempfile.gettempdir(), 'schedule_debug.html')
with contextlib.suppress(Exception), open(debug_path, 'w', encoding='utf-8') as f:
f.write(table_html)
data = _parse_schedule_html(table_html)
logger.info(list(data['weekdays'].keys()))
logger.info(f'Schedule parsed: {len(data["weekdays"])} days, classes={data["classes"]}')
return data
except Exception as e:
logger.error(f'Error parsing schedule: {e}')
table_html = await page.evaluate("""
() => {
const t = document.querySelector('table.schedule-table');
return t ? t.outerHTML : null;
}
""")
if not table_html:
logger.error('table.schedule-table not found in DOM')
return None
finally:
with contextlib.suppress(Exception):
await asyncio.wait_for(page.close(), timeout=5)
with contextlib.suppress(Exception):
await asyncio.wait_for(context_browser.close(), timeout=5)
debug_path = os.path.join(tempfile.gettempdir(), 'schedule_debug.html')
with contextlib.suppress(Exception), open(debug_path, 'w', encoding='utf-8') as f:
f.write(table_html)
data = _parse_schedule_html(table_html)
logger.info(list(data['weekdays'].keys()))
logger.info(f'Schedule parsed: {len(data["weekdays"])} days, classes={data["classes"]}')
return data
except Exception as e:
logger.error(f'Error parsing schedule: {e}')
return None
finally:
with contextlib.suppress(Exception):
await asyncio.wait_for(browser.close(), timeout=5)
except TimeoutError:
logger.error('Schedule fetch timed out (60s)')
return None
await page.close()
await context_browser.close()
finally:
await browser.close()
except Exception as e:
logger.error(f'Playwright error in schedule fetch: {e}')
return None
@@ -1602,13 +1485,10 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
int: Number of webinars found, or None if check failed
"""
logger.info('Running webinar check...')
_t0 = time.time()
_metric_check_start()
phpsessid = redis_client.get(KEY_PHPSESSID)
if not phpsessid:
logger.warning('PHPSESSID missing. Skipping check.')
_metric_check_fail(time.time() - _t0, phpsessid_missing=True)
# --- DEBUG LOGGING ---
try:
with open('phpsessid_missing.log', 'a') as f:
@@ -1622,88 +1502,78 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
content = ''
try:
async with asyncio.timeout(90):
async with async_playwright() as p:
# Connect to remote Playwright service
browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15)
async with async_playwright() as p:
# Connect to remote Playwright service
browser = await p.chromium.connect(PLAYWRIGHT_WS)
try:
# Create browser context with user agent
context_browser = await browser.new_context(user_agent=USER_AGENT)
# Add PHPSESSID cookie
await context_browser.add_cookies(
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}]
)
# Create new page
page = await context_browser.new_page()
try:
# Create browser context with user agent
context_browser = await browser.new_context(user_agent=USER_AGENT)
# Navigate to webinar page
await page.goto(WEBINAR_URL, wait_until='domcontentloaded')
# Add PHPSESSID cookie
await context_browser.add_cookies(
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}]
)
# Wait for the table to load
await page.wait_for_selector('#meetings table', timeout=10000)
await page.wait_for_timeout(2000)
# Create new page
page = await context_browser.new_page()
# Get page content
content = await page.content()
try:
# Navigate to webinar page
await asyncio.wait_for(page.goto(WEBINAR_URL, wait_until='domcontentloaded'), timeout=30)
# Check if "no webinar" message is present
if NO_WEBINAR_MARKER not in content:
logger.info('!!! WEBINAR FOUND !!!')
# Wait for the table to load
await page.wait_for_selector('#meetings table', timeout=10000)
await page.wait_for_timeout(2000)
# Extract webinar details from table rows
rows = page.locator('#meetings table tbody tr')
count = await rows.count()
# Get page content
content = await page.content()
for i in range(count):
row = rows.nth(i)
text = await row.inner_text()
# Check if "no webinar" message is present
if NO_WEBINAR_MARKER not in content:
logger.info('!!! WEBINAR FOUND !!!')
if NO_WEBINAR_MARKER not in text:
# Extract name (topic) from first column
name_elem = row.locator('td').nth(0)
name = await name_elem.inner_text()
name = name.strip()
# Extract webinar details from table rows
rows = page.locator('#meetings table tbody tr')
count = await rows.count()
# Extract join URL from fourth column
url_elem = row.locator('td').nth(3).locator('a[href*="/webinar/join/"]').first
url = await url_elem.get_attribute('href')
for i in range(count):
row = rows.nth(i)
text = await row.inner_text()
if name and url:
current_webinars.append({'name': name, 'url': url, 'text': text.strip()})
logger.info(f'Found webinar: {name} -> {url}')
else:
logger.info('No webinars found (expected message present)')
if NO_WEBINAR_MARKER not in text:
# Extract name (topic) from first column
name_elem = row.locator('td').nth(0)
name = await name_elem.inner_text()
name = name.strip()
# Extract join URL from fourth column
url_elem = row.locator('td').nth(3).locator('a[href*="/webinar/join/"]').first
url = await url_elem.get_attribute('href')
if name and url:
current_webinars.append({'name': name, 'url': url, 'text': text.strip()})
logger.info(f'Found webinar: {name} -> {url}')
else:
logger.info('No webinars found (expected message present)')
except Exception as e:
logger.error(f'Error checking page: {e}. Saving content for debug.')
# If page content is available, save it on error
with contextlib.suppress(Exception):
if page and not content:
content = await page.content()
_metric_check_fail(time.time() - _t0)
return None
finally:
with contextlib.suppress(Exception):
await asyncio.wait_for(page.close(), timeout=5)
with contextlib.suppress(Exception):
await asyncio.wait_for(context_browser.close(), timeout=5)
finally:
except Exception as e:
logger.error(f'Error checking page: {e}. Saving content for debug.')
# If page content is available, save it on error
with contextlib.suppress(Exception):
await asyncio.wait_for(browser.close(), timeout=5)
if page and not content:
content = await page.content()
return None
finally:
await page.close()
await context_browser.close()
finally:
await browser.close()
except TimeoutError:
logger.error('Webinar check timed out after 90s (playwright hang)')
_metric_check_fail(time.time() - _t0)
return None
except Exception as e:
logger.error(f'Playwright error: {e}')
_metric_check_fail(time.time() - _t0)
return None
# --- DEBUG LOGGING (Saving last response content) ---
@@ -1767,7 +1637,6 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
else:
logger.info(f'Found {len(current_webinars)} webinar(s), but all are already known')
_metric_check_ok(time.time() - _t0)
return len(current_webinars)
@@ -1811,8 +1680,6 @@ def main():
job_queue = app.job_queue
job_queue.run_repeating(check_webinars_job, interval=WEBINAR_CHECK_INTERVAL, first=10)
start_metrics_server()
logger.info('Bot started polling...')
app.run_polling()
-21
View File
@@ -1,21 +0,0 @@
# Error pages
Static HTTP error pages served by an Nginx image built in CI.
Edit the HTML in `html/`; the Dockerfile copies it into the image.
Kubernetes exposes `error-pages-service` in `error-pages` for Traefik's error
middleware. Keep the middleware's namespace and port aligned with that Service.
For a local build, run `docker build -t homelab-error-pages .` from this directory.
Compose references the private registry image rather than a build context.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n error-pages
kubectl get events -n error-pages --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,7 +1,7 @@
services:
errorpage:
build: .
image: gcr.forust.xyz/forust/error-pages:prod
image: gcr.forust.xyz/forust/error-pages:latest
pull_policy: build
container_name: error-pages
restart: unless-stopped
+1 -17
View File
@@ -20,8 +20,6 @@ spec:
selector:
matchLabels:
app: error-pages
strategy:
type: Recreate
template:
metadata:
labels:
@@ -29,21 +27,7 @@ spec:
spec:
containers:
- name: error-pages
image: gcr.forust.xyz/forust/error-pages:prod
# p95 6M, max 10M, no limit before.
resources:
requests:
cpu: "10m"
memory: "32Mi"
limits:
memory: "128Mi"
image: gcr.forust.xyz/forust/error-pages:latest
ports:
- containerPort: 80
readinessProbe:
httpGet:
path: /404.html
port: 80
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
---
-25
View File
@@ -1,25 +0,0 @@
# Gitea
Git hosting with HTTP and a separate SSH route.
Kubernetes uses the shared PostgreSQL service and `gitea-pvc` for repositories
and application data. Match the Gitea database password with the shared database
Secret. SSH is routed through Traefik's TCP entrypoint on 2221.
Compose uses a separate PostgreSQL 14 database, bind mounts `gitea-data/` and
`gitea-db/`, and publishes host port 2221. It is an alternative deployment with
its own database, not a second frontend for the Kubernetes instance.
Back up repositories, application configuration, and a consistent database dump
together. Gitea Actions definitions for this repository live in `../.gitea/`.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n gitea
kubectl get events -n gitea --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+3 -6
View File
@@ -1,6 +1,6 @@
services:
server:
image: docker.gitea.com/gitea:28.0.0
image: docker.gitea.com/gitea:1.27.3
container_name: gitea
restart: always
environment:
@@ -13,12 +13,9 @@ services:
- GITEA__database__PASSWD=gitea
- GITEA__database__NAME=gitea
# Server
- GITEA__server__ROOT_URL=https://git.forust.xyz
- GITEA__server__ROOT_URL=https://gitea.forust.xyz
- GITEA__server__SSH_DOMAIN=gitssh.forust.xyz
- GITEA__server__SSH_PORT=2221
# Pin 28.0 defaults explicitly (see k8s/config.yaml for rationale)
- GITEA__service__DISABLE_REGISTRATION=true
- GITEA__actions__RUN_RETENTION_DAYS=90
# Mailer
- GITEA__mailer__ENABLED=true
- GITEA__mailer__FROM=${SERVICE_EMAIL}
@@ -37,7 +34,7 @@ services:
- "traefik.http.services.gitea.loadbalancer.server.port=3000"
# Prod Router
- "traefik.http.routers.gitea.rule=Host(`git.forust.xyz`) || Host(`gitea.forust.xyz`)"
- "traefik.http.routers.gitea.rule=Host(`gitea.forust.xyz`)"
- "traefik.http.routers.gitea.entrypoints=websecure"
- "traefik.http.routers.gitea.tls.certresolver"
# Local Router
-30
View File
@@ -1,30 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: gitea-prod-tls
namespace: gitea
spec:
secretName: gitea-prod-tls
dnsNames:
- gcr.forust.xyz
- gitea.forust.xyz
- git.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: gitea
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+2 -11
View File
@@ -4,14 +4,11 @@ metadata:
name: gitea-config
namespace: gitea
data:
GITEA__server__ROOT_URL: "https://git.forust.xyz"
GITEA__server__DOMAIN: "gitea.forust.xyz"
GITEA__server__ROOT_URL: "https://gitea.forust.xyz"
GITEA__server__SSH_DOMAIN: "gitssh.forust.xyz"
GITEA__server__SSH_PORT: "2221"
GITEA__service__DISABLE_REGISTRATION: "true"
GITEA__actions__RUN_RETENTION_DAYS: "90"
GITEA__database__DB_TYPE: "postgres"
GITEA__database__HOST: "postgres.database.svc.cluster.local:5432"
GITEA__database__NAME: "gitea"
@@ -20,12 +17,6 @@ data:
GITEA__mailer__ENABLED: "false"
# No code/issue search needed: bleve reindexes the whole issue index on
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
# the rotational disk for an hour. "db" serves issue search from postgres.
GITEA__indexer__ISSUE_INDEXER_TYPE: "db"
GITEA__indexer__REPO_INDEXER_ENABLED: "false"
GITEA__log__logger.access.MODE: "console, file"
USER_UID: "1000"
USER_GID: "1000"
+4 -6
View File
@@ -24,8 +24,6 @@ spec:
selector:
matchLabels:
app: gitea
strategy:
type: Recreate
template:
metadata:
labels:
@@ -33,7 +31,7 @@ spec:
spec:
containers:
- name: gitea
image: gitea/gitea:28.0.0
image: docker.gitea.com/gitea:1.27.3
envFrom:
- configMapRef:
name: gitea-config
@@ -49,10 +47,10 @@ spec:
mountPath: /data
resources:
requests:
memory: "320Mi"
cpu: "100m"
memory: "512Mi"
cpu: "300m"
limits:
memory: "1Gi"
memory: "1.5Gi"
cpu: "1300m"
volumes:
- name: gitea-data
+9 -6
View File
@@ -7,18 +7,24 @@ spec:
entryPoints:
- websecure
routes:
- match: Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)
- match: Host(`gitea.forust.xyz`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: gitea-service
port: 3000
- match: Host(`gcr.forust.xyz`) && PathPrefix(`/v2`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: gitea-service
port: 3000
tls:
secretName: gitea-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -29,7 +35,7 @@ spec:
entryPoints:
- websecure
routes:
- match: (Host(`gitea.workstation.internal`) || Host(`gitea.gigaforust.internal`)) || (Host(`git.workstation.internal`) || Host(`git.gigaforust.internal`))
- match: Host(`gitea.workstation.internal`) || Host(`gitea.gigaforust.internal`)
kind: Rule
services:
- name: gitea-service
@@ -39,9 +45,6 @@ spec:
services:
- name: gitea-service
port: 3000
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRouteTCP
-26
View File
@@ -1,26 +0,0 @@
# Glance
Dashboard pages for links, service checks, and Docker containers.
Compose mounts `config/` and `assets/`. The Kubernetes equivalents are embedded
in `k8s/glance-config.yaml`: `glance-config` holds pages and `glance-assets` holds
`user.css`. Update both copies when changing shared content.
Kubernetes serves the dashboard under `/glance`. Its pod also mounts the node's
Docker socket. It references `glance-secrets` for `ADGUARD_PASSWORD`, but there is
no tracked Secret example; create that Secret in `glance` before starting it.
Compose expects a local `.env` with the same password.
The CSS mount points at the wrong ConfigMap on the reviewed main commit;
`fix/glance-assets` corrects it.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n glance
kubectl get events -n glance --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
File renamed without changes.
-29
View File
@@ -1,29 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: glance-prod-tls
namespace: glance
spec:
secretName: glance-prod-tls
dnsNames:
- forust.xyz
- www.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: glance
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+2 -2
View File
@@ -177,8 +177,8 @@ data:
# url: https://gitssh.forust.xyz
# - title: gcr.forust.xyz
# url: https://gcr.forust.xyz/v2/
- title: git.forust.xyz
url: https://git.forust.xyz
- title: gitea.forust.xyz
url: https://gitea.forust.xyz
- title: nextcloud.forust.xyz
url: https://nextcloud.forust.xyz
- title: mc.forust.xyz
+2 -4
View File
@@ -20,8 +20,6 @@ spec:
selector:
matchLabels:
app: glance
strategy:
type: Recreate
template:
metadata:
labels:
@@ -59,10 +57,10 @@ spec:
resources:
requests:
cpu: "50m"
memory: "32Mi"
memory: "64Mi"
limits:
cpu: "200m"
memory: "128Mi"
memory: "256Mi"
volumes:
- name: glance-config
configMap:
+1 -4
View File
@@ -16,7 +16,7 @@ spec:
- name: glance-service
port: 8080
tls:
secretName: glance-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -35,9 +35,6 @@ spec:
services:
- name: glance-service
port: 8080
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: Middleware
-18
View File
@@ -1,18 +0,0 @@
# Headscale
Headscale, Headplane, and a separate web administration UI on Docker.
Kubernetes only provides routes to the Docker host. Update the addresses in
`k8s/routing/external-service.yaml` if the host moves.
Copy `config/headscale.yaml.example`, `config/headplane.yaml.example`, and
`config/policy.json.example` to their names without `.example`. Set the public
server URL, DNS settings, Headplane cookie secret, and Headscale public URL.
The example URLs are placeholders.
Compose publishes Headscale on 18080, its metrics port on 19090, Headplane on
13000, and the other UI on 10080. The data volumes store the Headscale database,
keys, and Headplane state. The embedded DERP configuration needs reachable
addresses; Compose does not publish its UDP 3478 listener.
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,6 +1,6 @@
services:
headscale:
image: headscale/headscale:v0.29.4
image: headscale/headscale:0.29.3
restart: unless-stopped
container_name: headscale-server
command: serve
-41
View File
@@ -1,41 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: headscale-prod-tls
namespace: headscale
spec:
secretName: headscale-prod-tls
dnsNames:
- hs.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: headplane-prod-tls
namespace: headscale
spec:
secretName: headplane-prod-tls
dnsNames:
- hp.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: headscale
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+10 -7
View File
@@ -23,16 +23,22 @@ spec:
port: 8080
- match: Host(`hs.forust.xyz`) && PathPrefix(`/admin`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: headscale-ui-external
port: 80
- match: Host(`hs.forust.xyz`) && PathPrefix(`/metrics`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: headscale-server-external
port: 9090
tls:
secretName: headscale-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -47,6 +53,8 @@ spec:
kind: Rule
middlewares:
- name: headplane-prefix
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: headplane-external
port: 3000
@@ -56,7 +64,7 @@ spec:
- name: headplane-external
port: 3000
tls:
secretName: headplane-prod-tls
certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -82,9 +90,6 @@ spec:
services:
- name: headscale-server-external
port: 9090
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -105,5 +110,3 @@ spec:
services:
- name: headplane-external
port: 3000
tls:
secretName: internal-wildcard-tls
-4
View File
@@ -1,4 +0,0 @@
SECRET_ENCRYPTION_KEY="REPLACE_ME"
TZ="Europe/Bratislava"
PUID="1000"
PGID="1000"
-24
View File
@@ -1,24 +0,0 @@
# Homarr
Dashboard with Kubernetes integration and persistent application state.
Kubernetes uses the `homarr` ServiceAccount and the read-only ClusterRole in
`k8s/rbac.yaml`. Application data lives in `homarr-pvc`; supply the encryption key
from `k8s/secrets.yaml.example` before the first start and retain it with backups.
The committed ingress is internal. There is no `k8s/active` marker even though
manifests exist, so the workflow does not select Homarr automatically.
Compose publishes ports 80 and 81, mounts appdata and the Docker socket, and
expects a local kubeconfig. Check these host ports against Traefik before use.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n homarr
kubectl get events -n homarr --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
Loaded 100 of 433 files, more files were not shown because too many files have changed in this diff. Show more