Every bouncer-protected request blocks on a synchronous
GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403
when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi
on a single replica, so idle lookups measured 1.3-7.4s and a deploy
burst pushed them past the fork's implicit 10s default: gitea answered
403 for ten seconds straight, and containerd turned those 403s on
gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled
pods.
Three changes, plus the 403 feedback loop that made it sticky:
* gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route.
A deploy fires hundreds of parallel authenticated OCI requests
(runner Action API, manifest inspect per own image, containerd pulls,
smoke probes) and scanners gain nothing from a registry that already
does its own token auth. The web UI route keeps the bouncer.
* crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on
purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs),
so a second replica would corrupt the decision store.
* crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the
implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang.
LePresidente/http-generic-403-bf then banned us for our own 403s: five
POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT
address 192.168.88.1 that the Gitea Actions runner presents to Traefik
is not covered by the home-dynamic-IP whitelist. That scenario cannot
be dropped per-scenario - it is baked into the hub item
crowdsecurity/http-generic-bf v0.9, and disabling the whole
base-http-scenarios collection would cost ~40 useful detections. So the
janitor now deletes its decisions hourly and a new forust/lan
whitelist postoverflow covers 192.168.88.0/24. First janitor run
removed 165 decisions; none of the remaining ones are local.
Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way
parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth
challenge, LAPI at 60m CPU with no throttling.
registry_digest was written to return an empty string for a ref the registry
does not have, so render_pinned could print "cannot resolve <ref>" and stop.
It could not do that. Every caller runs under set -euo pipefail, pipefail
reports the rightmost non-zero stage, and the failed docker manifest inspect
made the assignment itself fail, which set -e turns into an immediate exit.
render_pinned therefore died silently on the first unresolvable ref: nothing on
stderr, nothing on stdout, exit 1. The apply loop piped that empty stream into
kubectl, so the whole deploy stopped with "error: no objects passed to apply" -
kubectl guessing at a cause, with the actual reason nowhere in the log. The
missing message is the reason the f54589a run looked like a network death.
Reproduced against the old file with a docker stub that always fails: identical
to the ac0f845 log. The || true makes the empty string reachable, and the apply
loop now names the file that failed instead of letting kubectl speak.
USERBOT_IMAGE named userbot:latest, but the build stopped pushing latest when
images became :prod. render_pinned resolves every gcr.forust.xyz reference it
finds in a manifest, not just the ones it rewrites, so that env var made the
whole apply abort with "cannot resolve userbot:latest; applying nothing".
Naming it :prod also makes the build job rebuild userbot, which is what
restores forust/userbot:prod - the tag is listed in the registry but its
manifest 404s, so the image: field in userbots.yaml had nothing to resolve
either.
Found by resolving every digest apply-k8s will need before spending a run on
it: 6 of 8 resolved, and both failures were userbot.
The first deploy to actually run died on its very first action, and the
error the other job reported was only the consequence.
DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy
is unprivileged, and this Arch host has no /var/backups at all, so
snapshot_dir's mkdir -p had to create it under root-owned /var and got
Permission denied. It refused to go on, which is exactly what the guard
is for, so no workload was touched - but the verify job then found no
pointer and could only say to go look by hand.
Defaulting to the deploy user's own XDG state directory fixes it with no
root and no setup step, and keeps the guard: an unwritable snapshot dir
still stops the deploy before the first apply.
ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is
overridable without editing the library. Verified on the workstation as
the unprivileged user: pointer published, commit recorded, 71 workload
generations and three helm releases captured, and the stale-pointer
refusal still works.
The pinned CI tools and the pre-push verification script were living in
a temp directory that did not survive. tmp/ is now ignored, so ruff
(which respects .gitignore by default) and the git ls-files globs the
lint jobs use both skip it without needing to be told.
The runner executes jobs on the host rather than in a container, so a
workflow that says 'python3' or 'npm' is really saying 'whatever this
machine happens to have today'. Both of the test jobs added in c00a472
were red on the first run for exactly that reason.
uv venv with no --python takes the first interpreter it finds, which is
the host's 3.14 here. pyrogram's sync.py calls the bare
asyncio.get_event_loop() that 3.14 no longer auto-creates, so three
tests died at collection. Pinned to 3.13, which is both what uv will
fetch when the host has none and what python:3.13-slim actually builds.
npm was missing outright, and turned up an hour later as npm 12 on node
26 - the same push, minutes apart. Neither is the panel image's
node:22-alpine, so node is now installed from the official tarball the
way the other tools are, at the image's major. npm is checked by running
it rather than by looking it up, so a name that resolves to something
broken reads as not installed.
Renovate keeps NODE_VERSION in step with the Dockerfile's node: tag, and
the two are one grouped dependency: CI that tests on a different major
than it builds on is a gate that can pass over a real break.
verify-k8s watches rollouts, which reports that pods converged. It cannot tell
a converged pod from a serving one. A Service selector pointing at a port
nothing listens on, a 500 from the app itself, a Traefik route that stopped
matching, a pod that OOMKilled early enough to still count as Available for the
duration of the check -- all of those are green at the rollout level and broken
for whoever opens the URL.
So ask what users ask. A smoke stage probes the public route of every active
service and fails on a transport error, a 5xx, or a 000, which curl reports
when it exits cleanly and nothing replied. Everything else passes, including 4xx:
a 404 from a path the service does not serve and a 302 to a login both prove
Traefik matched the host, the Service resolved to a pod and the pod answered,
which is the whole claim being tested.
An empty host list is an error, not a pass. Zero names means the extraction
broke, and reporting a clean deploy off a broken grep is the failure mode this
job exists to catch.
It runs on always() and after verify-k8s rather than before it, because a
rollback is when a route most needs re-checking. It only skips when verify-k8s
did, which is when nothing was deployed at all.
Two things worth writing down, because both were wrong on the first pass:
Stripping comments before reading the routes is not optional. naio and xui are
still in the tree commented out, and a plain grep picks both up and then reports
two services as unreachable when nobody ever deployed them. The apex
forust.xyz also needs a filter that admits it, so /\.forust\.xyz$/ quietly
dropped the site root.
And the 5xx test was written as ${code%%[0-9]*} != 5, which is empty for every
three-digit code, so a 500 was reported as ok. A case glob on the leading digit
is what actually works.
23 routes answer today, in 1.6s. The 5xx branch is the one part no live service
here exercises, so it was checked by running the block over 200 through 599 and
000 rather than against a real response.
The local bentopdf router matched Host(`pdf.wokstation.internal`), a hostname
that resolves to nothing. converters/k8s next to it has always had the correct
spelling, so the service is reachable in the path that actually deploys and this
one is dormant -- converters/ has no active marker, so select_manifests skips
the stack. It would only bite whoever switches the service back to Compose and
then wonders why one of the three internal routes 404s.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two things, both about not finding out late.
No workflow declared `permissions`, so all eighteen jobs across the four
workflows ran on a token with the default full repository scope. Every one of
them only checks out code, and deploy reaches the cluster over SSH with the
deploy key, and Renovate writes through its own bot PAT rather than the
Actions token. So `contents: read` is all any of them needed.
The panel image ships 15 known advisories and nothing was looking. Add a
scan-deps job that fails on anything new, and record the eight current ones by
ID in the workflow. It is a list rather than a baseline count so that the diff
that accepts an advisory says so in words, and it lives in our workflow
instead of the package manifest so a subtree sync from forust/userbot cannot
quietly widen the exemption.
Both halves were checked to fail on a regression, not just to pass today:
removing one --ignore-vuln turns the Python step red, and dropping
--audit-level to moderate turns the npm one red on the devalue advisory.
npm audits production dependencies only. All seven findings in the full tree
are build- or test-time: the esbuild advisory needs a vite dev server exposed
to the internet, and nanoid's infinite loop needs a custom generator called
with size 0, which postcss does not do. None are in the 91 kB bundle the panel
serves, so failing on them would be noise that trains people to ignore the
job.
The starlette entries are the reason the job is not "fail on everything":
fastapi 0.115.12 pins starlette<0.47.0 and the last four fixes need 0.49.1
through 1.3.1, so clearing them is a jump to fastapi 0.141.x and is upstream's
call, not a drive-by. Four of the seven are reachable in principle, which the
comment on the job sets out. The panel answers only on
userbot.workstation.internal with no public route, which is what keeps those
four from being an internet-facing DoS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
prettier, ruff, yamllint and hadolint were the only CI tools still called bare,
straight off whatever the runner happened to have installed. Pin them in
tool-versions.env like the other three and install them the same way, so the
versions Renovate moves are the versions CI runs.
Each pinned version equals what is already on the runner, so this changes what
CI does not at all today. It changes what CI does on a rebuilt runner: the
pinned one gets installed over the drift.
The four need four different mechanisms, which is why this is not one pattern:
hadolint a bare binary per platform, like actionlint
ruff,
yamllint PyPI wheels, unpacked by uv
prettier an npm tarball, unpacked by tar
prettier is the awkward one. Its entry point requires ../package.json relative
to its own real path, so copying the single file out -- which is what every
other installer here does -- yields a module-not-found at the first run. It
keeps its package directory in a versioned one next to a relative symlink, and
the tarball ships bin/ without the exec bit, so that needs chmod too.
hadolint's release names one platform uname-style and the other Go-style
(x86_64 but arm64), which 404s on the first architecture if you assume
otherwise.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
svelte-check sits in devDependencies at ^4.7.3 and nothing in the repository
ever invoked it, so the type errors it reports had no path to a human. Point a
script at it and run it in the frontend job, and it is clean: 0 errors, 0
warnings.
It shares the one `npm ci` with the test step. A second install would have
doubled the slowest part of the job to learn exactly the same thing.
The lockfile is untouched, because scripts are not part of what it pins.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The panel ships 25 pytest tests and 2 vitest tests. Nothing executed them:
there was no job, no local dev loop, and nothing that would have noticed when
one of them rotted. They pass, and they are 8 seconds of work, which is the
argument for having them.
Both jobs mirror how the image is built rather than how a developer would run
them by hand: `npm ci` because that is what the Dockerfile does, so the tree
under test is the tree that ships, and requirements-dev.txt through uv, which
is now pinned like the other CI tools.
The backend job runs `python -m pytest`, not bare `pytest`. The tests import
`app.*` relative to the backend directory, and only the `-m` form puts the
working directory on sys.path.
ruff format --check joins ruff check in the lint job. It needed 0ae0df7 to be
addable, since ten files disagreed with the style ruff.toml has always
declared.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The panel backend has 25 pytest tests that no workflow has ever run. Making
them run needs a throwaway virtualenv, and uv is what builds it in seconds
against the pinned requirements-dev.txt.
Installing it through install-ci-tools.sh rather than assuming it is on the
runner keeps the version in one place, where the other three tools already
live, and where the Renovate regex manager can move it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ruff.toml has declared `quote-style = "single"` and line-length 120 since the
lint job landed, and 118 of 128 files follow it. The panel backend and the
netbox configuration were written in black/prettier style instead, so a
`ruff format --check` would have failed on them from the start.
Bring them onto the style the repository already declares, which is what makes
the check adoptable at all. Formatting only: apart from quote style the diff is
multi-line expressions joined where they fit inside 120 columns.
Both suites still pass afterwards (25 pytest, and ruff check is clean).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`kubectl apply --dry-run=server` persists nothing, but it does execute the
admission webhooks of the real API server. The validate job runs on
pull_request with no branch guard, so anyone able to open a PR could run
arbitrary manifest content through cert-manager and Traefik in production.
Limit the step to pushes to main. A pull request loses nothing by it: only
main is ever deployed, and this job has to complete successfully before the
deploy workflow is allowed to start, so a bad CRD is still caught before
anything reaches the cluster -- on the push instead of on the PR.
The skip is announced rather than silent, so a missing server-side pass does
not read as a pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in 1505b63 therefore could not recover from a bad image.
Pin the digest at apply time. A digest is not knowable when a manifest is
written, so render_pinned resolves it on the way into the cluster and the
digest is never committed. Git keeps a readable `:prod`, Renovate keeps
seeing exactly the manifests it saw before, and the previous revision of
each workload now holds the digest that was actually serving, so undo
restores those exact bytes.
imagePullPolicy is dropped from the manifests rather than set to
IfNotPresent: a reference that is not `:latest` already defaults to it, and
that is what the Kubernetes docs ask for alongside a digest.
An unresolvable image is fatal instead of a warning, because carrying on
would quietly apply a mutable tag again.
restart_stale_images keeps its comparison but is no longer how a rebuild
reaches the cluster -- the pinned template rolls out on its own now. What
is left is a drift check for hand-run `kubectl set image`, so it matches
the container by repository: a pod's status now reports `repo@sha256:...`
while the manifest still says `:prod`.
The build job stops pushing `:latest` altogether, which removes the tag
that a dev branch could otherwise move under a prod deploy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Six of eight workloads had no readinessProbe, so a pod turned Ready the
moment its process started. The verify job relies on `rollout status`, so
it passed for images that crash-looped or served errors, which left the
rollback safety net inert.
Each probe targets the path the service is actually reached on:
- homepages: / (verified 200)
- error-pages: /404.html, the path Traefik's errorPages middleware
requests. / returns 403 by design and would never pass.
- webinar-checker: /health (verified 200). /metrics also answers, but it
is a Prometheus endpoint, not a readiness signal.
The two userbot deployments stay without probes: they expose no port and
no session file, and the panel reaches Telegram through its own client. A
truthful signal there needs a health endpoint in the app itself.
Reloader has been running since 23 September and is what makes the
reloader.stakater.com/auto annotation on a pod template do anything, but the
repository only held a namespace. It was a release someone installed by hand,
so it was invisible to review, invisible to Renovate, and one reinstall away from
being silently dropped.
Declaring it in HELM_RELEASES pins the chart version somewhere Renovate can
update it, and the active marker means the namespace is applied before the
upgrade instead of only existing as a side effect of the original install.
The values file sets nothing the running release does not already do, apart from
resource requests and limits, which the chart leaves empty.
Our manifests pin images to `:latest`, so a rebuild leaves the pod template
byte-identical. kubectl apply sees no change, creates no ReplicaSet and pulls
nothing, and the cluster keeps serving the previous build. imagePullPolicy:
Always does not help, because it only decides whether a pod that *is* starting
pulls, and no pod ever starts.
All eight workloads that consume an image from our own registry were affected.
Three of them had been running code from 23 September, and the single hardcoded
`rollout restart deployment/userbot-panel` covered one of the eight.
Restarting everything unconditionally was not the answer either: that bounces
healthy services on every deploy, error-pages included, and the brief window
where nothing answers is exactly what error-pages exists to prevent. So compare
what each workload actually runs against what the tag resolves to now, and
restart only the ones that differ. When the tag still points at the running
digest nothing happens, so a redeploy that changed no image is a no-op.
Scope is the repository, deliberately. Five more workloads run our images but
have no manifest here, and they are applied out of band. Walking the manifests
rather than the cluster means this can never reach them.
The digest is resolved for the node architecture. A multi-arch tag also carries
`unknown/unknown` entries for the build attestation, and a pod's imageID is
always the per-platform digest, so comparing the wrong entry would mark
everything stale forever.
Once a restart happens it bumps the generation, which is what makes the change
visible to changed_workloads and therefore watchable and revertible by the
verify stage.
verify_workloads ended on `[ -s "$failed_file" ]`, which is the opposite
of what its own contract says. A non-empty file means something failed, so
the function returned success exactly when a workload never came up, and
failure when everything was fine. Every rollback was therefore skipped,
and every deploy that changed anything ended red with an empty failure
list and a bogus "Rolled back successfully".
Worse, the check only ever ran at the end of stage_apply_k8s, inside the
same process as the apply. A job killed by timeout-minutes, cancelled by
a new push, or cut off by a dropped SSH connection never reached it, which
is precisely when a rollback matters. The three helm upgrades alone can
consume the whole 30-minute job budget, so that path was reachable.
Verification now lives in its own job, gated on always(), so it runs
whatever happened to the apply. The apply stage publishes its pre-apply
snapshot through DEPLOY_SNAPSHOT_DIR/current before touching anything,
and the verify stage picks it up from there. A snapshot whose recorded
commit does not match the deploy is refused rather than trusted, so a
stale pointer from an earlier run cannot make the rollback revert the
wrong workloads. An unwritable snapshot directory now fails the deploy up
front instead of silently continuing without a way back.
cancel-in-progress becomes false for the same reason: cancelling a run
kills the apply job and takes the verify job with it, which is the failure
this change exists to prevent. Both applies are idempotent, so queueing
costs little. The SSH key moves to a per-run directory removed on exit,
and the deploy is pinned to the exact commit CI validated.
The config lived in renovate.json at the repo root while everything else
Renovate-related sat under renovate/, and renovate/config.js was a second,
unused source of truth. Both are gone: renovate/renovate.json is now the
only config file.
Because the CronJob in the cluster cannot read the repository, its
ConfigMap carries an inlined copy of the config. That copy is generated,
and sync-renovate-configmap.sh --check now fails the build when it drifts
from the source file.
The workflows also stop carrying a copy of the renovate/renovate image
tag. They read it from renovate/k8s/cronjob.yaml, so the version validated
in CI is the version that actually runs in the cluster.
ci.yaml validates the config with renovate-config-validator, checks the
generated ConfigMap, and kubeconforms the CronJob's own manifests.
Adds three lint jobs (actionlint, shellcheck, compose) and a server-side
dry-run of the active manifests. Previously the only k8s check was
kubeconform, which has no schemas for CRDs, so every IngressRoute,
Certificate, PrometheusRule and Middleware was silently skipped.
The server-side pass needs the live API server because that is the only
place the real CRD schemas and the cert-manager / Traefik admission
webhooks exist. It is scoped to services carrying a k8s/active marker,
since dry-run needs the target namespace to exist. userbot/ is excluded
from shellcheck: it is a git subtree, and linting upstream's scripts would
let a routine subtree pull turn the deploy gate red on code we do not own.
kubeconform, shellcheck and actionlint are now installed from pinned
versions in tool-versions.env rather than picked up from the runner's
PATH. The Compose helper is shared with the deploy workflow so both
check the same file set the same way.
python -c 'exec /usr/bin/curl ...' is shell syntax, not Python, so every probe raised SyntaxError and the pod never became ready. Verified curl against /login/ returns HTTP 200.
checker.py initialises last_success to 0, so right after a pod restart
`time() - last_success` equals the current epoch. The rule compared that
against 300, went firing instantly, and humanizeDuration rendered the raw
epoch delta as ~20722d. The last_run > 0 guard did not help because a run
happens long before the first success.
Guard the duration rule on last_success > 0, keeping the duration
expression on the left of `and` so $value stays the real gap, and add a
separate WebinarCheckerNeverSucceeded rule for the zeroed-gauge case so a
checker that has never succeeded is still caught.
Queued deploy runs were deploying origin/main at start anyway (preflight reset), so waiting runs duplicated the newest deploy instead of their own commit. Cancel them.
compose.yaml is a local-only stand on 127.0.0.1:8000 per README; the live service runs in-cluster. The root active marker made apply-compose fail on gitignored .env vars.
2026.09.13-d4ce87c23 was never published (upstream tags month without leading zero); rollout stuck in ImagePullBackOff. Verified replacement tag exists on Docker Hub.
netbird runs in-cluster; the root active marker made apply-compose pick up netbird/compose.yaml and fail on gitignored .env vars. Drop the compose marker and enable k8s/active instead.
k0s runs kube-controller-manager, kube-scheduler and etcd inside its own
process rather than as pods, so the chart's Services never get endpoints
and the targets stay permanently absent. Drop the matching ServiceMonitors
and their Down/HighCommitDurations rules; kube-proxy and kubelet do get
endpoints on k0s and stay enabled.
docker compose config required real values for gitignored .env files and secrets, so validate always failed on stacks with :? guards (netbird, netbox). Validate structure only via --no-interpolate, --no-env-resolution and --no-path-resolution, keeping normalization and consistency checks.
Deploy workflow uses git-tracked manifests, DISABLED flag and kustomize overlays; add webinar-checker metrics with ServiceMonitor and alerts; upgrade shared postgres to 17 with statuspage DB and probes/resources.
All prod IngressRoutes switch tls.certResolver to tls.secretName
backed by per-router Certificates (HTTP-01, letsencrypt-prod).
adguard-prod reuses the shared adguard-certs secret (also feeds
DoT :853); sync CronJob removed as redundant.
Traefik certificatesResolvers removed: its internal
acme-http@internal router hijacks HTTP-01 for every host while
enabled, blocking external solvers. Dormant files (kener,
downtify) converted for consistency but not applied; n8n
untouched per live-only rule.
- CronJob mirrors Traefik prod cert dns.forust.xyz into
adguard-certs (cert-manager HTTP-01 is hijacked by Traefik
acme-http router, see adguardhome/k8s/cert-sync.yaml)
- reloader for auto-restart on secret rotation
- drop retired adguard.forust.xyz from prod route
- traefik: enable kubernetesIngress, drop unused staging resolver
Collection verbs cannot combine with resourceNames (grant would
be void). Instance verbs stay name-scoped to adguard-deployment;
list/watch is namespace-scoped (single Deployment in ns).
- grant list+watch on adguard-deployment (rollout status hung
without it, job hit activeDeadline and failed)
- compare content digests only (old hash embedded filenames, so
every run patched + restarted even when in sync)
Keeps explicit rollout restart alongside reloader annotation:
one extra restart per rotation (~60d) is accepted for
determinism if reloader is down.
adguard.forust.xyz is NXDOMAIN (host retired); the combo SAN cert kept
failing and dns.forust.xyz served TRAEFIK DEFAULT CERT. Scope prod route
to dns.forust.xyz only.
Also commit live traefik-values state (remove letsencrypt-staging
resolver, live since helm rev 33).
CronJob adguard-cert-sync (daily 03:17) copies the public cert/key for
dns.forust.xyz from Traefik acme.json into Secret adguard-certs, which
AdGuard mounts for DNS-over-TLS on :853.
- least-privilege RBAC: read pods/exec in ns traefik, get/update/patch
Secret adguard-certs and get/patch adguard-deployment in ns adguard
- script selects the PROD resolver entry only, matches main domain or
SANs, compares sha256 hashes, patches the secret and restarts the
deployment ONLY on change; exits non-zero and touches nothing when
Traefik holds no cert yet (HTTP-01 currently cannot complete)
VLESS+WS inbound (port 10000) via Traefik IngressRoute, panel on internal domains only with public route commented out. gitignore now covers nested k8s secrets and local-only grafana values.
Ruff and yamllint ship in Arch repos, drop their container pulls. Prettier keeps the node container but mounts persistent npm cache. Hadolint and kubeconform stay containerized (AUR-only / hermetic pin).
Drop duplicated .github/workflows (helm blocks ported to .gitea deploy first). Add ci concurrency with cancel on branches, pin checkout to SHA and prettier to 3.9.8, registry layer cache for image builds. renovate-run gains config validation and dry-run input.
renovate-run workflow_dispatch runs pinned renovate via docker on self-hosted runner. Sync helm-values manager into config.js/configmap, bump compose and validator pins to 44.97.2.
Move the shared postgres service from 15.19 to 17.6 as the postgres17 StatefulSet with its own PVC, extend the initdb and ingress policy with the statuspage database, and drop the now-unused per-app postgres manifests for authentik, gitea and netronome.
Move the shared postgres service from 15.19 to 17.6 as the postgres17 StatefulSet with its own PVC, extend the initdb and ingress policy with the statuspage database, and drop the now-unused per-app postgres manifests for authentik, gitea and netronome.
Protect public Traefik routes with CrowdSec HTTP decisions and restore access logging for web traffic analysis.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Keep the existing Redis PVC and data while migrating the edu-master workload from Deployment to StatefulSet.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
_collect_event_times re-clicked every a.event-link and waited ~3s per
event for a visible span.data, but the calendar embeds all times in
div.event-full-info[data-event-full-info-id] span.data already. The old
loop took ~109s for 25 events and collected 0 (original divs stay
sf-hidden), effectively hanging /diary. Now a single evaluate reads all
times (~3.7s), parsing HH:MM from p.date span.data.
- fetch each event's time via Playwright (click event-link, read span.data, close fancybox) and render as 'title (HH:MM)'
- '08:00' placeholder renders as localized 'unknown' (time_unknown key in ru/uk/en)
- diary week view shows Mon-Fri only (title ends at Friday)
- diary month view skips Sat/Sun by weekday_idx with name-based fallback
- schedule keyboard drops Sat/Sun day buttons
- restore-seed-job.yaml.example: translate runbook to English
- secrets.yaml.example: translate section headers to English
- webinar-checker.yaml: translate initContainer dependency-order comments to English
Co-authored-by: assistant
- ensure_prerequisites runs on startup, not per-request; kube config
errors surface as 503 PanelError
- serialize provisioning with a lock; drop per-endpoint prereq checks
- guard SPA fallback against path traversal (relative_to)
- add backend tests for auth flow, k8s service, spa routing; ci comment
for legacy userbot deployments
services marked k8s/active are applied via kubectl; the rest via docker
compose. inactive services with k8s/ keep only routing manifests
(external Services, EndpointSlices, Ingresses) to reach docker backends.
headscale/nextcloud routing moved to k8s/routing/.
validations: compose config --quiet + kubectl apply --dry-run=client.
namespace manifests applied first. pull_policy:build stacks get
build+push before up so the registry image stays fresh.
- rename ambiguous loop var, merge nested if (E741, SIM102)
- use tempfile.gettempdir() for debug dump (S108)
- drop unused total_lessons assignment (F841)
- fix parser: capture tr attrs via finditer, strip HTML comments before
cell parse (was leaving '-->' in subject names)
- store class choice in redis: user:{id}:schedule_class (private) and
chat:{id}:schedule_class (groups, admin-only via /setclass)
- /schedule renders day for stored class, /setclass sets it directly
- drop teacher emoji, format grade as "N клас"
- Create .venv with pyrofork and all dependencies installed
- Add pyrightconfig.json at workspace root and in userbot/
- Enable pyright as language server for Python in Zed settings
- Update uv.lock with resolved dependency tree
- S113: Add timeout=10 to all requests calls (74 fixes)
- E722: Replace bare except: with except Exception:
- B904: Replace redundant re-raise with bare raise
- E402: Add noqa for intentional late imports after import_library()
- S102/S307/S310/S311/S603/S605/S606/S607/S108: Add noqa for intentional usage
- F601: Fix duplicate dict key in unsplash.py
- N802: Rename ReplyCheck -> reply_check with backward compat alias
- N813: Rename bs -> BS in icons.py
- B007/B020: Rename loop var _j in animations.py
- SIM102: Collapse nested if in autofwd.py
- SIM113: Use enumerate() in calculator.py
- A002: Add noqa for builtin shadowing in admlist.py
- F811: Add noqa for cohere redefinition
- edu_master: Fix ARG001, S108, S110, SIM117, apply --unsafe-fixes
- Add modules_list.txt with full module inventory
- S113: Add timeout=10 to all requests calls (74 fixes)
- E722: Replace bare except: with except Exception:
- B904: Replace redundant re-raise with bare raise
- E402: Add noqa for intentional late imports after import_library()
- S102/S307/S310/S311/S603/S605/S606/S607/S108: Add noqa for intentional usage
- F601: Fix duplicate dict key in unsplash.py
- N802: Rename ReplyCheck -> reply_check with backward compat alias
- N813: Rename bs -> BS in icons.py
- B007/B020: Rename loop var _j in animations.py
- SIM102: Collapse nested if in autofwd.py
- SIM113: Use enumerate() in calculator.py
- A002: Add noqa for builtin shadowing in admlist.py
- F811: Add noqa for cohere redefinition
- edu_master: Fix ARG001, S108, S110, SIM117, apply --unsafe-fixes
- Add modules_list.txt with full module inventory
- Multi-stage build: pip deps built in separate stage
- Static ffmpeg binary instead of apt package (avoid 200+ deps)
- Keep only git, mediainfo, wget via apt
- Multi-stage build: pip deps built in separate stage
- Static ffmpeg binary instead of apt package (avoid 200+ deps)
- Keep only git, mediainfo, wget via apt
- Parse school diary calendar HTML table for daily/weekly/monthly views
- Cache diary data with 5-minute TTL
- Inline keyboard for today/tomorrow/week/month selection
- Ukrainian month/weekday names and formatting
| `nfs` | nfs-subdir-external-provisioner или nfs.csi.k8s.io | Retain | Данные, которые нужно сохранять |
**Главное преимущество:** PVC c `storageClassName: nfs` при удалении namespace сохраняют данные на диске, так как NFS-провизор использует `reclaimPolicy: Retain` или файлы физически остаются в NFS-экспорте.
## 5. Ответы на частые вопросы
**В:** Не упадёт ли local-path при установке NFS?
**О:** Нет, они независимы. local-path продолжает работать как обычно.
**В:** Данные NFS и local-path будут на одном диске?
**О:** Да, можно настроить оба на одном разделе, в разных каталогах.
**В:** Что если у меня несколько нод?
**О:** NFS сервер нужно поднять на одной ноде, а с других нод должна быть доступна шари. Для multi-node лучше использовать отдельный сервер или distributed storage (Longhorn, Rook/Ceph).
**В:** Можно ли использовать существующий NTFS-раздел для NFS?
**О:** Не рекомендуется — NTFS не поддерживает права Linux (no_root_squash не сработает корректно), возможны проблемы с блокировками и производительностью.
and (webinar_check_last_success_timestamp_seconds > 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for:2m
labels:
severity:critical
annotations:
summary:"Webinar checker has no successful check for 5m"
description:"edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Checks are running but none has ever succeeded since pod start.
# Split out from the rule above so a zeroed gauge never feeds
and (webinar_check_last_run_timestamp_seconds > 0)
for:10m
labels:
severity:critical
annotations:
summary:"Webinar checker has never completed a successful check"
description:'edu-master/webinar-checker:checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki:{namespace="edu-master", container="webinar-checker"}).'
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
description:"edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
- alert:EduPhpsessidMissing
expr:|
edu_phpsessid_present == 0
for:10m
labels:
severity:critical
annotations:
summary:"EDU_PHPSESSID missing"
description:"edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker and playwright deployments unavailable.
description:"edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
RUN pip install --no-cache-dir requests==2.32.3 redis==5.2.1
# Copy application code
COPY . .
Loaded 100 of 480 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.