Renovate: group all minor updates and all patch updates into reviewable PRs. Crowdsec: whitelist static VPS IP. Remove stale cloudflared/k8s/active marker.
Move L3 enforcement to the host firewall-bouncer (systemd, nftables): drop the Traefik plugin, its secrets volume and the crowdsec Middleware, remove bouncer refs from all IngressRoutes. Disable the http-generic-bf scenario (403-burst bans hurt legit automation under L3 enforcement). Add a Gateway API PoC for homepages prod and CrowdSec PrometheusRule alerts.
Single replica is the whole ingress; liveness kills at 2s timeouts took every public service down in a loop. Same stall-sized budgets as postgres/metallb. Applied live via helm (pinned 41.5.0, values from git).
apply-k8s and apply-compose share a workstation flock so host docker churn never overlaps cluster churn. Each helm upgrade and the apply loop wait up to 10m for load <28 first, so a deploy never piles onto an already-hot node (the load-40/netbird-death/pending-helm spiral).
The pin step writes manifest PUTs on every main push, but login was gated on services != ''. Manifest-only pushes skipped login and pushed anonymously (401); this only worked before via a stale persistent login on the old runner.
Deployment was stuck in ImagePullBackOff and only generated log churn. Workload, service, routes and certs removed from the cluster; PVC, config and manifests kept so it can be re-enabled by restoring k8s/active.
Callers prepend BIN_DIR to PATH only after the script exits, so the bare uv invocation in install_uv_tool died with 127 on clean runners. Export BIN_DIR to PATH inside the script and invoke the just-installed binary by absolute path.
Pushes no longer reach the cluster by default; set the AUTODEPLOY repo variable to 'true' to re-enable, or dispatch manually. Skips propagate through the existing needs chain.
Set AUTODEPLOY=false under Settings -> Actions -> Variables and pushes stop deploying with no commit; unset means on. Manual Run workflow always bypasses the switch. The existing needs/skipped chaining propagates the skip through verify and smoke untouched.
Postmaster was SIGKILLed in a loop: 70s fsync stalls on the loaded rotational disk outlasted the 5-minute startup budget and the 60s liveness tolerance, and every kill bought another full WAL replay. Startup budget 15min, liveness 5x60s. Already applied live with kubectl; this keeps git in sync.
The service label the rule grouped by is the scrape target name, not the backend: with honorLabels=false the per-route value traefik emits is renamed to exported_service, so the old rule measured one global aggregate and both exclusions matched nothing. Group the latency and 5xx alerts by exported_service with exclusions in traefik's real namespace-routename-hash format. Already applied live with kubectl; this commit keeps git in sync.
No grafana is deployed from this stack, so grafana-prod/grafana-local only ever served 503s. Live objects already deleted directly; this keeps git from resurrecting them on the next deploy.
Reverts the v16 revert: data migrated via pg_dumpall from PG14 into a fresh PG16 data directory (same extension versions vectorchord 0.4.3/pgvectors 0.2.0). Verified 1944 assets, 2 users, 10 albums on 16.10. Note: the restore OOMed once at the 2Gi limit during index builds and recovered via WAL replay; worth watching under PG16 load.
userbot/ (bot plus panel) now lives at /home/forust/userbot as a clone of forust/userbot instead of a subtree in homelab. Cleans up the pipeline references that only existed for it: scan-deps, test-backend and test-frontend jobs, the userbot build matrix entries, the deploy panel hook, and the userbot-only pyrightconfig. Running cluster workloads are untouched; homelab just stops building and testing upstream's code.
SignalExchange ConnectStream holds 60s gRPC streams by design; at night they exceed 5% of samples and pin P95 to the 5.0s bucket ceiling, flapping the alert. Also fixes the stale xui exclusion pattern, which matched no real service label.
A single blink of the registry failed render_pinned for the whole file and redded the apply stage. registry_digest now retries 3 times under a 25s timeout with a warning per attempt; empty still means unresolvable and callers report it by name as before.
RollingUpdate with default maxSurge needs a spare pod the single node does not have (99% CPU requested), so multi-workload restarts end Pending and verify times out. Recreate on all replicas:1 Deployments (immich-server and bentopdf keep RollingUpdate at replicas 2). Also trims CPU/memory requests toward measured use (adguard, authentik, gitea, netbox, uptime-kuma, netbird-server) and gives traefik requests/limits so it is no longer BestEffort.
Renovate branches only carry version/digest bumps, so scan-deps, test-backend, test-frontend and build just burn runner time on the box that also serves prod. Static checks and validate still run. Digest and patch updates automerge (playwright, helm and major rules below still override to no-automerge). Also replaces deprecated helm --atomic with --wait --rollback-on-failure.
cron.rebuild_issue_indexer runs at start, so every gitea pod restart reindexed the whole issue index and read ~7MB/s off the rotational disk for an hour.
The OOMKilled container behind the immich crash loop was invisible: PodCrashLooping only fires once kubelet has already given up and started the backoff.
Mount traefik-plugins PVC at /plugins-storage instead of the chart default emptyDir, so the crowdsec-bouncer download survives node reboots. Without this Traefik starts before the network is ready, the download from plugins.traefik.io times out, plugins get disabled and every route behind the middleware returns 404/503 until a manual restart.
pg_isready with the 1s default times out under I/O stall and kubelet kills a healthy postgres mid-recovery; each kill restarts a multi-minute fsync from zero and loops forever. readiness/liveness timeout 5s, liveness threshold 5, startup budget 15min.
An --atomic upgrade whose own rollback never finishes leaves the release in pending-*, blocking every future run until a human rolls back (loki rev 18/21). Recover automatically before and after each upgrade, and fail loud when recovery does not land on deployed. Also skip helm-managed workloads in rollback_workloads: rollout undo there would step back to the revision --atomic just escaped.
Chart default is required podAntiAffinity on hostname plus RollingUpdate 25%/25%, which is maxUnavailable=0 at replicas=1: the new pod stays Unschedulable while the old one lives, and the old one never leaves while the new one is not Ready. Null the affinity (an empty map deep-merges with the default and keeps the rule) and set maxSurge/maxUnavailable to 1.
Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.
Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.
Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.
prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.
Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.
CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.
Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
Portainer had been running for 111 days with a 512Mi request and a 2Gi limit
against 50M of measured use, on a node that is short of memory. It is a UI over
the Docker socket; nothing in the repo or the cluster depends on it.
The marker goes, not the manifests. `portainer/k8s/active` is what puts these
files in the deploy's manifest set, so without it the next push leaves the
namespace alone and the manifests stay on disk for a one-command return. This
also matters for the smoke stage: that host list is built from the active
directories, so `portainer.forust.xyz` leaves it and the new router check does
not go looking for a route to a service we just retired.
In the cluster the Deployment, the Service and both IngressRoutes are deleted.
The routes go first: leaving an IngressRoute behind a deleted Service keeps a
Traefik router pointing at nothing, which answers 502 while looking perfectly
healthy to the stage that just started checking for routers.
Deliberately kept, so this is reversible rather than destructive: the namespace,
the 2Gi `portainer-data-pvc` and both Certificates stay. Re-enabling is
`git checkout HEAD~1 -- portainer/k8s/active` plus an apply, and no Let's Encrypt
quota is spent reissuing the production certificate.
`glance` still links to `portainer.forust.xyz` and that tile will now be a 404.
Left alone on purpose. The Cloudflare record is manual and cfddns only ever
creates records, so `portainer.forust.xyz` keeps resolving until it is removed
in the dashboard, same as `dockmon.forust.xyz`.
Verified: no Traefik router matches portainer any more, all 19 hosts left in the
smoke list still have a router, and 16/16 local gates pass.
Six pods reserved far more memory than they have ever touched. uptime-kuma held
a 3Gi limit against 469M of measured p95, metube 2Gi against 72M, convertx
1.5Gi against 85M, netbird-server 1Gi against 97M, searxng 700Mi against 134M
and bentopdf 700Mi against 4M. Every one of them is a ceiling the scheduler
counts against the node while the memory sits unused.
Requests move down with the limits but never below the measured p95, so none of
these becomes an eviction candidate as a side effect of being right-sized. The
limits keep between 2.2x and 11.6x over the observed max, which is the figure
that decides whether a pod gets OOM-killed during a burst.
Net effect across the six: requests -557M, limits -4.6Gi, all of it ceiling that
was never in use. This is the first change that actually gives memory back.
CPU limits are left exactly as they were. They were not part of the sizing pass,
they are not being hit on a node sitting at 5% CPU, and removing them is a
separate decision from moving memory.
Verified: each limit is above the container's own observed max and each request
is above its p95, and 16/16 local gates pass.
Eighteen containers had no memory limit at all, so nothing on the node could
bound them. Three of the values files even claimed to set resources: Helm does
not complain about a key it does not recognise, so the block sat there looking
like a limit while the pod ran unbounded.
alloy is the one that mattered. The chart reads `alloy.resources`; the file had
`controller.resources`, so the DaemonSet that tails every pod log on the node
shipped with nothing at all. `kubeStateMetrics` is the same trap in a different
shape -- that is the condition key, the values live under `kube-state-metrics` --
and `configReloader` in the alloy chart sits at the top level rather than under
`alloy`. Each one is verified by rendering the chart and reading the resources
back off the containers, because a values key that is ignored looks exactly
like one that works.
reloader turned out to be set and still wrong: 64Mi request against a measured
p95 of 73M, so the pod ran permanently above its own request and stayed a
standing eviction candidate. That is the pod that restarts every other pod, so
it is the last one that should be evicted. Raised to 96Mi.
Requests are set at p95 throughout, grafana, playwright and alloy included.
Left at the values first proposed they would have sat below their own p95 and
queued for eviction ahead of everything smaller. CPU limits are deliberately
absent: the node is I/O bound at 5% CPU, and CFS throttling would turn disk
wait into runnable-throttled, which is the failure mode that took the node down.
The prometheus and alertmanager configReloader sidecars are left open: chart
86.2.3 does not template the key, so reaching those two containers needs a
postRenderer.
Verified: all four charts render with the resources landing on the intended
containers, and 16/16 local gates pass.
The smoke stage treats any HTTP response as proof the service is serving,
which is right -- a 302 to a login or a 404 from a path the app does not
serve still means the chain is intact. But a 404 is not evidence of that on
its own: a router Traefik refused to build answers with exactly the same
404 and nothing behind it.
That is not hypothetical. The crowdsec bouncer is a plugin, and when Traefik
cannot fetch it at startup it disables the plugin without failing, then drops
every router whose chain referenced it. Sixteen routes answered 404 and the
stage printed `ok` for all sixteen, because a dropped router and an unserved
path are indistinguishable from outside.
The Kubernetes objects cannot tell us either: the IngressRoute is still
sitting there looking healthy, the router Traefik built from it is simply not
there. So ask Traefik. api.insecure is already on for the internal entrypoint
and the router list says which hosts it matches right now.
Every probed host has to appear in that list. HTTP routers only -- the TCP
ones match on a HostSNI wildcard and the UDP ones carry no rule at all, both
selected by entrypoint and port, so neither can answer the question. A router
mid-rollout is legitimately absent for a moment, so the list is re-read twice
over 20s; a plugin that failed to load stays absent and waiting cannot rescue
it. An unreadable router list fails the stage rather than skipping the check,
since a check that cannot run is not a passing check.
Verified against the live cluster: all 23 public routes have a router and the
stage passes. With gitea, grafana and uptime removed from that list the
probes still answer and the stage fails on exactly those three.
Every service tracked the mutable `:prod` tag, so a deploy applied whatever
that tag happened to name at the time rather than the commit it was
deploying. A rollback had no way to state what it was rolling back to, and
two deploys of one commit could land different images.
CI now publishes an immutable `sha-<commit12>` tag beside `:prod` on main,
and re-tags it for every image a push did not rebuild. That re-tag copies
the manifest list, so no layer moves. The deploy resolves the immutable tag
to a digest and pins the workload to it, and only falls back to the moving
tag when the immutable one cannot be resolved -- which it says out loud,
because that fallback is the deploy quietly ceasing to be reproducible from
its own commit.
The image list comes out of the tree with git grep rather than being written
out a second time, so adding a service no longer means keeping two lists in
step.
build also gains the three jobs it was skipping -- scan-deps, test-backend,
test-frontend -- so a change that breaks them cannot be tagged at all. The
two run blocks where a mid-loop failure was survivable now run under
set -euo pipefail: the build loop and the service detector both carried on
past an error and could report a green build having produced nothing.
The registry password moves from run: substitution into an env: block. A
quote, a backtick or a $(...) in the password is parsed as shell before the
command ever runs, and a login that failed that way looked exactly like a
build that failed.
The apply and verify timeouts stay at 45 and 30 minutes. The comments now
record the arithmetic that says so rather than leaving the numbers to be
raised on the next scare: three no-op helm upgrades run 3-5 minutes, one
broken release is a single 10 minute rollback because the loop aborts on
the first failure, and the apply loop itself is about a minute. That is
roughly 15 minutes of work against a 45 minute budget. verify is 32
workloads at 8 wide -- four waves of 300 seconds, 20 minutes -- which
leaves room for two serial rollbacks, and only becomes derivable at 45 once
rollback_workloads is parallelised.
The parser-stage whitelist already covers 84.245.64.0/18, so an event
from the phone is dropped before it reaches a bucket and no decision is
ever created for it - confirmed against 72h of traefik access logs, where
the phone shows up as 84.245.120.147, inside that /18. But that left the
bouncer's own ClientTrustedIPs without the range, so the guarantee rested
on a single config. If the parser whitelist ever stops matching, a ban
would be created and then served against the phone, which is the one
thing that must not happen: the address belongs to a carrier, so it comes
back to us by rotation and a 4h ban is not survivable from the device.
ClientTrustedIPs bypasses the bouncer and the decision cache entirely, so
repeating the range there holds even if a decision exists for any reason.
All nine parser-stage ranges are now mirrored in the bouncer, and the
bouncer has no range the parser stage does not know about.
The file header now records that this middleware must be applied together
with a traefik restart. Applying it alone wedges the plugin: in stream
mode handleStreamTicker runs over package-level globals that no
reconfiguration stops, so every route referencing the middleware answers
404 with 'invalid middleware crowdsec-crowdsec-bouncer@kubernetescrd'
until the pod is replaced. Re-applying the prior config does not recover
it and the config is not the cause - NewChecker is a plain net.ParseCIDR
and cannot fail on a valid range. That cost 21 routes down before the
restart requirement was found; recovery is a pod replace, ~35s.
Verified live: middleware applied, traefik restarted, 17 of 20 hosts
serving (the three exceptions are unchanged and unrelated - searxng
returns its own 429, checkmk is down with 503, and one host is
local-only), zero invalid-middleware errors, and the LAPI still shows
/v1/decisions/stream polls at the 15s interval.
A connection that died silently used to hang until the job timeout, and
the stage was never re-run. One flaky TCP session cost a whole
45-minute apply, and the symptom - a job that stops mid-output with no
error - is what made the last few deploy failures expensive to read.
ServerAliveInterval/CountMax cap how long a dead peer goes unnoticed at
~60s, ConnectTimeout caps setup. Only exit 255 - ssh's own transport
failures - is retried, up to three attempts with a growing gap. A stage
that fails on its own merits exits with the remote's status, so a real
failure surfaces its own log immediately instead of being repeated three
times over 45 minutes. The stages are declarative applies, so re-running
one that had already committed is harmless.
The stage environment now goes through `env` as separate argv entries
rather than one interpolated string, so nothing in REPO, DEPLOY_SHA or
DEPLOY_SNAPSHOT_DIR is re-split by the remote shell.
Verified against a stubbed ssh: clean run attempts once, a single
transport failure recovers on attempt 2 and exits 0, three failures give
up preserving 255, and a stage failing with 1 or 7 attempts once and
passes the code through unchanged.
Also records why USERBOT_IMAGE stays on the prod tag: render_pinned
rewrites only plain `image:` lines, and this ref is what the panel
injects into the per-instance Deployments it creates, so those instances
track the tag rather than the panel's own resolved digest. The two
panel-created instances currently in the cluster are digest-pinned, so
the panel does accept one either way; the tag is the choice, not a
limitation.
dtek-notif is being picked up again, so leave its compose alone. It is
also the one image not rebuilt since the build dropped :latest, so
:prod does not exist for it yet - unlike the other six, which resolve.
Keeping this file out of the change also keeps dtek_notif/* out of the
build's changed-service detector, so pushing does not build an image
nobody asked for.
These five stacks still asked for :latest, but the build stopped pushing
it - on main it only pushes main and prod, on dev only dev. Every one of
these images therefore resolved only because the registry still had a
stale :latest from before that change, and the next time one of them was
actually built the reference would have dangled.
Named for dtek-notif: of the seven, only that one still resolved at
:prod, because it is the only image not rebuilt since the build dropped
:latest - and it is the only one of these five stacks the deploy does not
manage (no `active` marker, and its file is docker-compose.yaml, which
the COMPOSE_STACKS glob does not even match). Touching
dtek_notif/docker-compose.yaml matches dtek_notif/* in the build's
changed-service detector, so pushing this builds it and publishes
:prod for it too.
Verified the other six resolve at :prod in the registry.
Chasing why apply-k8s kept dying mid-run turned up a self-inflicted
ban loop. 585 of 586 LePresidente/http-generic-403-bf alerts in the LAPI
came from 193.181.211.79 - our own VPS - POSTing
/management.ManagementService/GetServerKey, i.e. the NetBird client's
own management call. netbird-server had never seen a single one of them,
so the 403 was not NetBird's: the bouncer was rejecting the request
before it got there. A banned peer keeps retrying, each retry is another
403, and the scenario turns five 403s in ten seconds into a 4h ban, so
the loop kept re-arming the ban it was serving. The hourly janitor step
that deleted those decisions hourly was masking all of it.
* netbird/k8s/ingress.yaml - drop the bouncer from the mesh API routes
(gRPC-gateway management, signal, relay, /api, /oauth2). A ban there
locks a peer out of the network it needs to reach anything else, and
those endpoints authenticate by NetBird token, not by a login form.
The dashboard keeps the bouncer; it is a real login surface.
netbird-local was already exempt, so this makes prod match.
* crowdsec-middleware.yaml - CrowdsecMode stream instead of live. live
blocked on GET /v1/decisions per request, so a burst saturated the
LAPI and the plugin 403'd IPs that were never banned. v1.3.3 ignores
UpdateMaxFailure in live, so fail-open is only reachable in stream;
-1 now means an unreachable LAPI degrades to "no protection" rather
than "every site 403". 15s poll instead of the 60s default, because
the runner shares one public IP with the house.
Also corrects the key name: HTTPTimeoutSeconds, not
CrowdsecLapiTimeout, which never existed and was being silently
dropped, leaving the 10s default. Back at 10, not the 2s f7cd75d
guessed - a pull that times out leaves the ban cache frozen at its
startup contents, so new bans would silently never apply.
* crowdsec-values.yaml - CIDR allowlisting moves to parsers/s02-enrich,
where CrowdSec's docs put it: a parser whitelist drops the event before
it reaches a bucket, so those addresses never become a decision at
all. The old postoverflow LAN list was checked only after the ban
existed, which is the window the deploy kept landing in. Added
100.64.0.0/10, which the RFC 1918 blocks miss and where the
workstation, the k0s node and the VPS actually live. The DDNS
home-IP whitelist stays in postoverflows, because resolving a hostname
is the expensive check the docs reserve that stage for.
ClientTrustedIPs mirrors that list so the bouncer skips the LAPI
round-trip entirely for those addresses.
* janitor-cronjob.yaml - drop step 5. The bouncer can no longer
manufacture 403s, so the only remaining firings of that scenario are
real scanners, and deleting their decisions hourly was undoing a
working ban.
* Also lands the LAPI config.yaml.local (SQLite WAL, Central API off,
bounded flush) that f7cd75d's comments referenced but never included:
the LAPI was blocked in fsync on its rollback journal, and the CAPI
resolver held a write transaction while timing out against a host
this network cannot reach.
Verified in-cluster: no new 403-bf alerts in the 3.5min after applying,
the management endpoint answers 404 from the backend instead of 403 from
the bouncer in 0.17s, and the VPS client reports Management and Signal
connected with 2/2 relays.
Every bouncer-protected request blocks on a synchronous
GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403
when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi
on a single replica, so idle lookups measured 1.3-7.4s and a deploy
burst pushed them past the fork's implicit 10s default: gitea answered
403 for ten seconds straight, and containerd turned those 403s on
gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled
pods.
Three changes, plus the 403 feedback loop that made it sticky:
* gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route.
A deploy fires hundreds of parallel authenticated OCI requests
(runner Action API, manifest inspect per own image, containerd pulls,
smoke probes) and scanners gain nothing from a registry that already
does its own token auth. The web UI route keeps the bouncer.
* crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on
purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs),
so a second replica would corrupt the decision store.
* crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the
implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang.
LePresidente/http-generic-403-bf then banned us for our own 403s: five
POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT
address 192.168.88.1 that the Gitea Actions runner presents to Traefik
is not covered by the home-dynamic-IP whitelist. That scenario cannot
be dropped per-scenario - it is baked into the hub item
crowdsecurity/http-generic-bf v0.9, and disabling the whole
base-http-scenarios collection would cost ~40 useful detections. So the
janitor now deletes its decisions hourly and a new forust/lan
whitelist postoverflow covers 192.168.88.0/24. First janitor run
removed 165 decisions; none of the remaining ones are local.
Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way
parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth
challenge, LAPI at 60m CPU with no throttling.
registry_digest was written to return an empty string for a ref the registry
does not have, so render_pinned could print "cannot resolve <ref>" and stop.
It could not do that. Every caller runs under set -euo pipefail, pipefail
reports the rightmost non-zero stage, and the failed docker manifest inspect
made the assignment itself fail, which set -e turns into an immediate exit.
render_pinned therefore died silently on the first unresolvable ref: nothing on
stderr, nothing on stdout, exit 1. The apply loop piped that empty stream into
kubectl, so the whole deploy stopped with "error: no objects passed to apply" -
kubectl guessing at a cause, with the actual reason nowhere in the log. The
missing message is the reason the f54589a run looked like a network death.
Reproduced against the old file with a docker stub that always fails: identical
to the ac0f845 log. The || true makes the empty string reachable, and the apply
loop now names the file that failed instead of letting kubectl speak.
USERBOT_IMAGE named userbot:latest, but the build stopped pushing latest when
images became :prod. render_pinned resolves every gcr.forust.xyz reference it
finds in a manifest, not just the ones it rewrites, so that env var made the
whole apply abort with "cannot resolve userbot:latest; applying nothing".
Naming it :prod also makes the build job rebuild userbot, which is what
restores forust/userbot:prod - the tag is listed in the registry but its
manifest 404s, so the image: field in userbots.yaml had nothing to resolve
either.
Found by resolving every digest apply-k8s will need before spending a run on
it: 6 of 8 resolved, and both failures were userbot.
The first deploy to actually run died on its very first action, and the
error the other job reported was only the consequence.
DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy
is unprivileged, and this Arch host has no /var/backups at all, so
snapshot_dir's mkdir -p had to create it under root-owned /var and got
Permission denied. It refused to go on, which is exactly what the guard
is for, so no workload was touched - but the verify job then found no
pointer and could only say to go look by hand.
Defaulting to the deploy user's own XDG state directory fixes it with no
root and no setup step, and keeps the guard: an unwritable snapshot dir
still stops the deploy before the first apply.
ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is
overridable without editing the library. Verified on the workstation as
the unprivileged user: pointer published, commit recorded, 71 workload
generations and three helm releases captured, and the stale-pointer
refusal still works.
The pinned CI tools and the pre-push verification script were living in
a temp directory that did not survive. tmp/ is now ignored, so ruff
(which respects .gitignore by default) and the git ls-files globs the
lint jobs use both skip it without needing to be told.
The runner executes jobs on the host rather than in a container, so a
workflow that says 'python3' or 'npm' is really saying 'whatever this
machine happens to have today'. Both of the test jobs added in c00a472
were red on the first run for exactly that reason.
uv venv with no --python takes the first interpreter it finds, which is
the host's 3.14 here. pyrogram's sync.py calls the bare
asyncio.get_event_loop() that 3.14 no longer auto-creates, so three
tests died at collection. Pinned to 3.13, which is both what uv will
fetch when the host has none and what python:3.13-slim actually builds.
npm was missing outright, and turned up an hour later as npm 12 on node
26 - the same push, minutes apart. Neither is the panel image's
node:22-alpine, so node is now installed from the official tarball the
way the other tools are, at the image's major. npm is checked by running
it rather than by looking it up, so a name that resolves to something
broken reads as not installed.
Renovate keeps NODE_VERSION in step with the Dockerfile's node: tag, and
the two are one grouped dependency: CI that tests on a different major
than it builds on is a gate that can pass over a real break.
verify-k8s watches rollouts, which reports that pods converged. It cannot tell
a converged pod from a serving one. A Service selector pointing at a port
nothing listens on, a 500 from the app itself, a Traefik route that stopped
matching, a pod that OOMKilled early enough to still count as Available for the
duration of the check -- all of those are green at the rollout level and broken
for whoever opens the URL.
So ask what users ask. A smoke stage probes the public route of every active
service and fails on a transport error, a 5xx, or a 000, which curl reports
when it exits cleanly and nothing replied. Everything else passes, including 4xx:
a 404 from a path the service does not serve and a 302 to a login both prove
Traefik matched the host, the Service resolved to a pod and the pod answered,
which is the whole claim being tested.
An empty host list is an error, not a pass. Zero names means the extraction
broke, and reporting a clean deploy off a broken grep is the failure mode this
job exists to catch.
It runs on always() and after verify-k8s rather than before it, because a
rollback is when a route most needs re-checking. It only skips when verify-k8s
did, which is when nothing was deployed at all.
Two things worth writing down, because both were wrong on the first pass:
Stripping comments before reading the routes is not optional. naio and xui are
still in the tree commented out, and a plain grep picks both up and then reports
two services as unreachable when nobody ever deployed them. The apex
forust.xyz also needs a filter that admits it, so /\.forust\.xyz$/ quietly
dropped the site root.
And the 5xx test was written as ${code%%[0-9]*} != 5, which is empty for every
three-digit code, so a 500 was reported as ok. A case glob on the leading digit
is what actually works.
23 routes answer today, in 1.6s. The 5xx branch is the one part no live service
here exercises, so it was checked by running the block over 200 through 599 and
000 rather than against a real response.
The local bentopdf router matched Host(`pdf.wokstation.internal`), a hostname
that resolves to nothing. converters/k8s next to it has always had the correct
spelling, so the service is reachable in the path that actually deploys and this
one is dormant -- converters/ has no active marker, so select_manifests skips
the stack. It would only bite whoever switches the service back to Compose and
then wonders why one of the three internal routes 404s.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two things, both about not finding out late.
No workflow declared `permissions`, so all eighteen jobs across the four
workflows ran on a token with the default full repository scope. Every one of
them only checks out code, and deploy reaches the cluster over SSH with the
deploy key, and Renovate writes through its own bot PAT rather than the
Actions token. So `contents: read` is all any of them needed.
The panel image ships 15 known advisories and nothing was looking. Add a
scan-deps job that fails on anything new, and record the eight current ones by
ID in the workflow. It is a list rather than a baseline count so that the diff
that accepts an advisory says so in words, and it lives in our workflow
instead of the package manifest so a subtree sync from forust/userbot cannot
quietly widen the exemption.
Both halves were checked to fail on a regression, not just to pass today:
removing one --ignore-vuln turns the Python step red, and dropping
--audit-level to moderate turns the npm one red on the devalue advisory.
npm audits production dependencies only. All seven findings in the full tree
are build- or test-time: the esbuild advisory needs a vite dev server exposed
to the internet, and nanoid's infinite loop needs a custom generator called
with size 0, which postcss does not do. None are in the 91 kB bundle the panel
serves, so failing on them would be noise that trains people to ignore the
job.
The starlette entries are the reason the job is not "fail on everything":
fastapi 0.115.12 pins starlette<0.47.0 and the last four fixes need 0.49.1
through 1.3.1, so clearing them is a jump to fastapi 0.141.x and is upstream's
call, not a drive-by. Four of the seven are reachable in principle, which the
comment on the job sets out. The panel answers only on
userbot.workstation.internal with no public route, which is what keeps those
four from being an internet-facing DoS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
prettier, ruff, yamllint and hadolint were the only CI tools still called bare,
straight off whatever the runner happened to have installed. Pin them in
tool-versions.env like the other three and install them the same way, so the
versions Renovate moves are the versions CI runs.
Each pinned version equals what is already on the runner, so this changes what
CI does not at all today. It changes what CI does on a rebuilt runner: the
pinned one gets installed over the drift.
The four need four different mechanisms, which is why this is not one pattern:
hadolint a bare binary per platform, like actionlint
ruff,
yamllint PyPI wheels, unpacked by uv
prettier an npm tarball, unpacked by tar
prettier is the awkward one. Its entry point requires ../package.json relative
to its own real path, so copying the single file out -- which is what every
other installer here does -- yields a module-not-found at the first run. It
keeps its package directory in a versioned one next to a relative symlink, and
the tarball ships bin/ without the exec bit, so that needs chmod too.
hadolint's release names one platform uname-style and the other Go-style
(x86_64 but arm64), which 404s on the first architecture if you assume
otherwise.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
svelte-check sits in devDependencies at ^4.7.3 and nothing in the repository
ever invoked it, so the type errors it reports had no path to a human. Point a
script at it and run it in the frontend job, and it is clean: 0 errors, 0
warnings.
It shares the one `npm ci` with the test step. A second install would have
doubled the slowest part of the job to learn exactly the same thing.
The lockfile is untouched, because scripts are not part of what it pins.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The panel ships 25 pytest tests and 2 vitest tests. Nothing executed them:
there was no job, no local dev loop, and nothing that would have noticed when
one of them rotted. They pass, and they are 8 seconds of work, which is the
argument for having them.
Both jobs mirror how the image is built rather than how a developer would run
them by hand: `npm ci` because that is what the Dockerfile does, so the tree
under test is the tree that ships, and requirements-dev.txt through uv, which
is now pinned like the other CI tools.
The backend job runs `python -m pytest`, not bare `pytest`. The tests import
`app.*` relative to the backend directory, and only the `-m` form puts the
working directory on sys.path.
ruff format --check joins ruff check in the lint job. It needed 0ae0df7 to be
addable, since ten files disagreed with the style ruff.toml has always
declared.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The panel backend has 25 pytest tests that no workflow has ever run. Making
them run needs a throwaway virtualenv, and uv is what builds it in seconds
against the pinned requirements-dev.txt.
Installing it through install-ci-tools.sh rather than assuming it is on the
runner keeps the version in one place, where the other three tools already
live, and where the Renovate regex manager can move it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ruff.toml has declared `quote-style = "single"` and line-length 120 since the
lint job landed, and 118 of 128 files follow it. The panel backend and the
netbox configuration were written in black/prettier style instead, so a
`ruff format --check` would have failed on them from the start.
Bring them onto the style the repository already declares, which is what makes
the check adoptable at all. Formatting only: apart from quote style the diff is
multi-line expressions joined where they fit inside 120 columns.
Both suites still pass afterwards (25 pytest, and ruff check is clean).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`kubectl apply --dry-run=server` persists nothing, but it does execute the
admission webhooks of the real API server. The validate job runs on
pull_request with no branch guard, so anyone able to open a PR could run
arbitrary manifest content through cert-manager and Traefik in production.
Limit the step to pushes to main. A pull request loses nothing by it: only
main is ever deployed, and this job has to complete successfully before the
deploy workflow is allowed to start, so a bad CRD is still caught before
anything reaches the cluster -- on the push instead of on the PR.
The skip is announced rather than silent, so a missing server-side pass does
not read as a pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in 1505b63 therefore could not recover from a bad image.
Pin the digest at apply time. A digest is not knowable when a manifest is
written, so render_pinned resolves it on the way into the cluster and the
digest is never committed. Git keeps a readable `:prod`, Renovate keeps
seeing exactly the manifests it saw before, and the previous revision of
each workload now holds the digest that was actually serving, so undo
restores those exact bytes.
imagePullPolicy is dropped from the manifests rather than set to
IfNotPresent: a reference that is not `:latest` already defaults to it, and
that is what the Kubernetes docs ask for alongside a digest.
An unresolvable image is fatal instead of a warning, because carrying on
would quietly apply a mutable tag again.
restart_stale_images keeps its comparison but is no longer how a rebuild
reaches the cluster -- the pinned template rolls out on its own now. What
is left is a drift check for hand-run `kubectl set image`, so it matches
the container by repository: a pod's status now reports `repo@sha256:...`
while the manifest still says `:prod`.
The build job stops pushing `:latest` altogether, which removes the tag
that a dev branch could otherwise move under a prod deploy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Six of eight workloads had no readinessProbe, so a pod turned Ready the
moment its process started. The verify job relies on `rollout status`, so
it passed for images that crash-looped or served errors, which left the
rollback safety net inert.
Each probe targets the path the service is actually reached on:
- homepages: / (verified 200)
- error-pages: /404.html, the path Traefik's errorPages middleware
requests. / returns 403 by design and would never pass.
- webinar-checker: /health (verified 200). /metrics also answers, but it
is a Prometheus endpoint, not a readiness signal.
The two userbot deployments stay without probes: they expose no port and
no session file, and the panel reaches Telegram through its own client. A
truthful signal there needs a health endpoint in the app itself.
Reloader has been running since 23 September and is what makes the
reloader.stakater.com/auto annotation on a pod template do anything, but the
repository only held a namespace. It was a release someone installed by hand,
so it was invisible to review, invisible to Renovate, and one reinstall away from
being silently dropped.
Declaring it in HELM_RELEASES pins the chart version somewhere Renovate can
update it, and the active marker means the namespace is applied before the
upgrade instead of only existing as a side effect of the original install.
The values file sets nothing the running release does not already do, apart from
resource requests and limits, which the chart leaves empty.
Our manifests pin images to `:latest`, so a rebuild leaves the pod template
byte-identical. kubectl apply sees no change, creates no ReplicaSet and pulls
nothing, and the cluster keeps serving the previous build. imagePullPolicy:
Always does not help, because it only decides whether a pod that *is* starting
pulls, and no pod ever starts.
All eight workloads that consume an image from our own registry were affected.
Three of them had been running code from 23 September, and the single hardcoded
`rollout restart deployment/userbot-panel` covered one of the eight.
Restarting everything unconditionally was not the answer either: that bounces
healthy services on every deploy, error-pages included, and the brief window
where nothing answers is exactly what error-pages exists to prevent. So compare
what each workload actually runs against what the tag resolves to now, and
restart only the ones that differ. When the tag still points at the running
digest nothing happens, so a redeploy that changed no image is a no-op.
Scope is the repository, deliberately. Five more workloads run our images but
have no manifest here, and they are applied out of band. Walking the manifests
rather than the cluster means this can never reach them.
The digest is resolved for the node architecture. A multi-arch tag also carries
`unknown/unknown` entries for the build attestation, and a pod's imageID is
always the per-platform digest, so comparing the wrong entry would mark
everything stale forever.
Once a restart happens it bumps the generation, which is what makes the change
visible to changed_workloads and therefore watchable and revertible by the
verify stage.
verify_workloads ended on `[ -s "$failed_file" ]`, which is the opposite
of what its own contract says. A non-empty file means something failed, so
the function returned success exactly when a workload never came up, and
failure when everything was fine. Every rollback was therefore skipped,
and every deploy that changed anything ended red with an empty failure
list and a bogus "Rolled back successfully".
Worse, the check only ever ran at the end of stage_apply_k8s, inside the
same process as the apply. A job killed by timeout-minutes, cancelled by
a new push, or cut off by a dropped SSH connection never reached it, which
is precisely when a rollback matters. The three helm upgrades alone can
consume the whole 30-minute job budget, so that path was reachable.
Verification now lives in its own job, gated on always(), so it runs
whatever happened to the apply. The apply stage publishes its pre-apply
snapshot through DEPLOY_SNAPSHOT_DIR/current before touching anything,
and the verify stage picks it up from there. A snapshot whose recorded
commit does not match the deploy is refused rather than trusted, so a
stale pointer from an earlier run cannot make the rollback revert the
wrong workloads. An unwritable snapshot directory now fails the deploy up
front instead of silently continuing without a way back.
cancel-in-progress becomes false for the same reason: cancelling a run
kills the apply job and takes the verify job with it, which is the failure
this change exists to prevent. Both applies are idempotent, so queueing
costs little. The SSH key moves to a per-run directory removed on exit,
and the deploy is pinned to the exact commit CI validated.
The config lived in renovate.json at the repo root while everything else
Renovate-related sat under renovate/, and renovate/config.js was a second,
unused source of truth. Both are gone: renovate/renovate.json is now the
only config file.
Because the CronJob in the cluster cannot read the repository, its
ConfigMap carries an inlined copy of the config. That copy is generated,
and sync-renovate-configmap.sh --check now fails the build when it drifts
from the source file.
The workflows also stop carrying a copy of the renovate/renovate image
tag. They read it from renovate/k8s/cronjob.yaml, so the version validated
in CI is the version that actually runs in the cluster.
ci.yaml validates the config with renovate-config-validator, checks the
generated ConfigMap, and kubeconforms the CronJob's own manifests.
Adds three lint jobs (actionlint, shellcheck, compose) and a server-side
dry-run of the active manifests. Previously the only k8s check was
kubeconform, which has no schemas for CRDs, so every IngressRoute,
Certificate, PrometheusRule and Middleware was silently skipped.
The server-side pass needs the live API server because that is the only
place the real CRD schemas and the cert-manager / Traefik admission
webhooks exist. It is scoped to services carrying a k8s/active marker,
since dry-run needs the target namespace to exist. userbot/ is excluded
from shellcheck: it is a git subtree, and linting upstream's scripts would
let a routine subtree pull turn the deploy gate red on code we do not own.
kubeconform, shellcheck and actionlint are now installed from pinned
versions in tool-versions.env rather than picked up from the runner's
PATH. The Compose helper is shared with the deploy workflow so both
check the same file set the same way.
python -c 'exec /usr/bin/curl ...' is shell syntax, not Python, so every probe raised SyntaxError and the pod never became ready. Verified curl against /login/ returns HTTP 200.
checker.py initialises last_success to 0, so right after a pod restart
`time() - last_success` equals the current epoch. The rule compared that
against 300, went firing instantly, and humanizeDuration rendered the raw
epoch delta as ~20722d. The last_run > 0 guard did not help because a run
happens long before the first success.
Guard the duration rule on last_success > 0, keeping the duration
expression on the left of `and` so $value stays the real gap, and add a
separate WebinarCheckerNeverSucceeded rule for the zeroed-gauge case so a
checker that has never succeeded is still caught.
Queued deploy runs were deploying origin/main at start anyway (preflight reset), so waiting runs duplicated the newest deploy instead of their own commit. Cancel them.
compose.yaml is a local-only stand on 127.0.0.1:8000 per README; the live service runs in-cluster. The root active marker made apply-compose fail on gitignored .env vars.
2026.09.13-d4ce87c23 was never published (upstream tags month without leading zero); rollout stuck in ImagePullBackOff. Verified replacement tag exists on Docker Hub.
netbird runs in-cluster; the root active marker made apply-compose pick up netbird/compose.yaml and fail on gitignored .env vars. Drop the compose marker and enable k8s/active instead.
k0s runs kube-controller-manager, kube-scheduler and etcd inside its own
process rather than as pods, so the chart's Services never get endpoints
and the targets stay permanently absent. Drop the matching ServiceMonitors
and their Down/HighCommitDurations rules; kube-proxy and kubelet do get
endpoints on k0s and stay enabled.
docker compose config required real values for gitignored .env files and secrets, so validate always failed on stacks with :? guards (netbird, netbox). Validate structure only via --no-interpolate, --no-env-resolution and --no-path-resolution, keeping normalization and consistency checks.
Deploy workflow uses git-tracked manifests, DISABLED flag and kustomize overlays; add webinar-checker metrics with ServiceMonitor and alerts; upgrade shared postgres to 17 with statuspage DB and probes/resources.
All prod IngressRoutes switch tls.certResolver to tls.secretName
backed by per-router Certificates (HTTP-01, letsencrypt-prod).
adguard-prod reuses the shared adguard-certs secret (also feeds
DoT :853); sync CronJob removed as redundant.
Traefik certificatesResolvers removed: its internal
acme-http@internal router hijacks HTTP-01 for every host while
enabled, blocking external solvers. Dormant files (kener,
downtify) converted for consistency but not applied; n8n
untouched per live-only rule.
- CronJob mirrors Traefik prod cert dns.forust.xyz into
adguard-certs (cert-manager HTTP-01 is hijacked by Traefik
acme-http router, see adguardhome/k8s/cert-sync.yaml)
- reloader for auto-restart on secret rotation
- drop retired adguard.forust.xyz from prod route
- traefik: enable kubernetesIngress, drop unused staging resolver
Collection verbs cannot combine with resourceNames (grant would
be void). Instance verbs stay name-scoped to adguard-deployment;
list/watch is namespace-scoped (single Deployment in ns).
- grant list+watch on adguard-deployment (rollout status hung
without it, job hit activeDeadline and failed)
- compare content digests only (old hash embedded filenames, so
every run patched + restarted even when in sync)
Keeps explicit rollout restart alongside reloader annotation:
one extra restart per rotation (~60d) is accepted for
determinism if reloader is down.
adguard.forust.xyz is NXDOMAIN (host retired); the combo SAN cert kept
failing and dns.forust.xyz served TRAEFIK DEFAULT CERT. Scope prod route
to dns.forust.xyz only.
Also commit live traefik-values state (remove letsencrypt-staging
resolver, live since helm rev 33).
CronJob adguard-cert-sync (daily 03:17) copies the public cert/key for
dns.forust.xyz from Traefik acme.json into Secret adguard-certs, which
AdGuard mounts for DNS-over-TLS on :853.
- least-privilege RBAC: read pods/exec in ns traefik, get/update/patch
Secret adguard-certs and get/patch adguard-deployment in ns adguard
- script selects the PROD resolver entry only, matches main domain or
SANs, compares sha256 hashes, patches the secret and restarts the
deployment ONLY on change; exits non-zero and touches nothing when
Traefik holds no cert yet (HTTP-01 currently cannot complete)
VLESS+WS inbound (port 10000) via Traefik IngressRoute, panel on internal domains only with public route commented out. gitignore now covers nested k8s secrets and local-only grafana values.
Ruff and yamllint ship in Arch repos, drop their container pulls. Prettier keeps the node container but mounts persistent npm cache. Hadolint and kubeconform stay containerized (AUR-only / hermetic pin).
Drop duplicated .github/workflows (helm blocks ported to .gitea deploy first). Add ci concurrency with cancel on branches, pin checkout to SHA and prettier to 3.9.8, registry layer cache for image builds. renovate-run gains config validation and dry-run input.
renovate-run workflow_dispatch runs pinned renovate via docker on self-hosted runner. Sync helm-values manager into config.js/configmap, bump compose and validator pins to 44.97.2.
Move the shared postgres service from 15.19 to 17.6 as the postgres17 StatefulSet with its own PVC, extend the initdb and ingress policy with the statuspage database, and drop the now-unused per-app postgres manifests for authentik, gitea and netronome.
Move the shared postgres service from 15.19 to 17.6 as the postgres17 StatefulSet with its own PVC, extend the initdb and ingress policy with the statuspage database, and drop the now-unused per-app postgres manifests for authentik, gitea and netronome.
Protect public Traefik routes with CrowdSec HTTP decisions and restore access logging for web traffic analysis.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Keep the existing Redis PVC and data while migrating the edu-master workload from Deployment to StatefulSet.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
_collect_event_times re-clicked every a.event-link and waited ~3s per
event for a visible span.data, but the calendar embeds all times in
div.event-full-info[data-event-full-info-id] span.data already. The old
loop took ~109s for 25 events and collected 0 (original divs stay
sf-hidden), effectively hanging /diary. Now a single evaluate reads all
times (~3.7s), parsing HH:MM from p.date span.data.
- fetch each event's time via Playwright (click event-link, read span.data, close fancybox) and render as 'title (HH:MM)'
- '08:00' placeholder renders as localized 'unknown' (time_unknown key in ru/uk/en)
- diary week view shows Mon-Fri only (title ends at Friday)
- diary month view skips Sat/Sun by weekday_idx with name-based fallback
- schedule keyboard drops Sat/Sun day buttons
- restore-seed-job.yaml.example: translate runbook to English
- secrets.yaml.example: translate section headers to English
- webinar-checker.yaml: translate initContainer dependency-order comments to English
Co-authored-by: assistant
- ensure_prerequisites runs on startup, not per-request; kube config
errors surface as 503 PanelError
- serialize provisioning with a lock; drop per-endpoint prereq checks
- guard SPA fallback against path traversal (relative_to)
- add backend tests for auth flow, k8s service, spa routing; ci comment
for legacy userbot deployments
services marked k8s/active are applied via kubectl; the rest via docker
compose. inactive services with k8s/ keep only routing manifests
(external Services, EndpointSlices, Ingresses) to reach docker backends.
headscale/nextcloud routing moved to k8s/routing/.
validations: compose config --quiet + kubectl apply --dry-run=client.
namespace manifests applied first. pull_policy:build stacks get
build+push before up so the registry image stays fresh.
- rename ambiguous loop var, merge nested if (E741, SIM102)
- use tempfile.gettempdir() for debug dump (S108)
- drop unused total_lessons assignment (F841)
- fix parser: capture tr attrs via finditer, strip HTML comments before
cell parse (was leaving '-->' in subject names)
- store class choice in redis: user:{id}:schedule_class (private) and
chat:{id}:schedule_class (groups, admin-only via /setclass)
- /schedule renders day for stored class, /setclass sets it directly
- drop teacher emoji, format grade as "N клас"
- Create .venv with pyrofork and all dependencies installed
- Add pyrightconfig.json at workspace root and in userbot/
- Enable pyright as language server for Python in Zed settings
- Update uv.lock with resolved dependency tree
- S113: Add timeout=10 to all requests calls (74 fixes)
- E722: Replace bare except: with except Exception:
- B904: Replace redundant re-raise with bare raise
- E402: Add noqa for intentional late imports after import_library()
- S102/S307/S310/S311/S603/S605/S606/S607/S108: Add noqa for intentional usage
- F601: Fix duplicate dict key in unsplash.py
- N802: Rename ReplyCheck -> reply_check with backward compat alias
- N813: Rename bs -> BS in icons.py
- B007/B020: Rename loop var _j in animations.py
- SIM102: Collapse nested if in autofwd.py
- SIM113: Use enumerate() in calculator.py
- A002: Add noqa for builtin shadowing in admlist.py
- F811: Add noqa for cohere redefinition
- edu_master: Fix ARG001, S108, S110, SIM117, apply --unsafe-fixes
- Add modules_list.txt with full module inventory
- S113: Add timeout=10 to all requests calls (74 fixes)
- E722: Replace bare except: with except Exception:
- B904: Replace redundant re-raise with bare raise
- E402: Add noqa for intentional late imports after import_library()
- S102/S307/S310/S311/S603/S605/S606/S607/S108: Add noqa for intentional usage
- F601: Fix duplicate dict key in unsplash.py
- N802: Rename ReplyCheck -> reply_check with backward compat alias
- N813: Rename bs -> BS in icons.py
- B007/B020: Rename loop var _j in animations.py
- SIM102: Collapse nested if in autofwd.py
- SIM113: Use enumerate() in calculator.py
- A002: Add noqa for builtin shadowing in admlist.py
- F811: Add noqa for cohere redefinition
- edu_master: Fix ARG001, S108, S110, SIM117, apply --unsafe-fixes
- Add modules_list.txt with full module inventory
- Multi-stage build: pip deps built in separate stage
- Static ffmpeg binary instead of apt package (avoid 200+ deps)
- Keep only git, mediainfo, wget via apt
- Multi-stage build: pip deps built in separate stage
- Static ffmpeg binary instead of apt package (avoid 200+ deps)
- Keep only git, mediainfo, wget via apt
- Parse school diary calendar HTML table for daily/weekly/monthly views
- Cache diary data with 5-minute TTL
- Inline keyboard for today/tomorrow/week/month selection
- Ukrainian month/weekday names and formatting
| `nfs` | nfs-subdir-external-provisioner или nfs.csi.k8s.io | Retain | Данные, которые нужно сохранять |
**Главное преимущество:** PVC c `storageClassName: nfs` при удалении namespace сохраняют данные на диске, так как NFS-провизор использует `reclaimPolicy: Retain` или файлы физически остаются в NFS-экспорте.
## 5. Ответы на частые вопросы
**В:** Не упадёт ли local-path при установке NFS?
**О:** Нет, они независимы. local-path продолжает работать как обычно.
**В:** Данные NFS и local-path будут на одном диске?
**О:** Да, можно настроить оба на одном разделе, в разных каталогах.
**В:** Что если у меня несколько нод?
**О:** NFS сервер нужно поднять на одной ноде, а с других нод должна быть доступна шари. Для multi-node лучше использовать отдельный сервер или distributed storage (Longhorn, Rook/Ceph).
**В:** Можно ли использовать существующий NTFS-раздел для NFS?
**О:** Не рекомендуется — NTFS не поддерживает права Linux (no_root_squash не сработает корректно), возможны проблемы с блокировками и производительностью.
echo":: warning::ssh transport failed, retrying (${attempt}/3)"
sleep $((attempt *5))
fi
rc=0
# shellcheck disable=SC2029 # remote_cmd/ssh_opts expand on the client on purpose: they select the local ssh invocation, only the heredoc runs remotely.
and (webinar_check_last_success_timestamp_seconds > 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for:2m
labels:
severity:critical
annotations:
summary:"Webinar checker has no successful check for 5m"
description:"edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Checks are running but none has ever succeeded since pod start.
# Split out from the rule above so a zeroed gauge never feeds
and (webinar_check_last_run_timestamp_seconds > 0)
for:10m
labels:
severity:critical
annotations:
summary:"Webinar checker has never completed a successful check"
description:'edu-master/webinar-checker:checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki:{namespace="edu-master", container="webinar-checker"}).'
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
description:"edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
- alert:EduPhpsessidMissing
expr:|
edu_phpsessid_present == 0
for:10m
labels:
severity:critical
annotations:
summary:"EDU_PHPSESSID missing"
description:"edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker and playwright deployments unavailable.
description:"edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
METRICS_PORT: "8000"
# Database
REDIS_HOST: "redis"
REDIS_PORT: "6379"
TZ: "Europe/Kyiv"
Loaded 100 of 459 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.