Commit Graph
49 Commits
Author SHA1 Message Date
renovate-bot 4ddbbbd2ed chore(deps): update all minor updates
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 10s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 6s
ci / lint-prettier (push) Successful in 16s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 11s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 6s
ci / build (push) Skipped
ci / lint-compose (pull_request) Successful in 8s
ci / lint-actionlint (pull_request) Successful in 4s
ci / lint-shellcheck (pull_request) Successful in 7s
ci / lint-prettier (pull_request) Successful in 13s
ci / lint-ruff (pull_request) Successful in 6s
ci / lint-yaml (pull_request) Successful in 11s
ci / lint-dockerfiles (pull_request) Successful in 6s
ci / validate (pull_request) Successful in 9s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 10s
2026-10-06 10:19:13 +00:00
forust a6af69dca0 fix(k8s): Recreate singletons and trim requests for scheduler headroom
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-compose (push) Successful in 3s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 18s
ci / build (push) Successful in 1m30s
RollingUpdate with default maxSurge needs a spare pod the single node does not have (99% CPU requested), so multi-workload restarts end Pending and verify times out. Recreate on all replicas:1 Deployments (immich-server and bentopdf keep RollingUpdate at replicas 2). Also trims CPU/memory requests toward measured use (adguard, authentik, gitea, netbox, uptime-kuma, netbird-server) and gives traefik requests/limits so it is no longer BestEffort.
2026-09-28 22:36:46 +02:00
forust f9e4623ade fix(k8s): size the remaining workloads against measured use
Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.

Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.

Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.

prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.

Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.

CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.

Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
2026-09-28 10:38:19 +02:00
forust 16aaeb60c1 fix(k8s): set requests and limits on the pods that shipped with neither
Eighteen containers had no memory limit at all, so nothing on the node could
bound them. Three of the values files even claimed to set resources: Helm does
not complain about a key it does not recognise, so the block sat there looking
like a limit while the pod ran unbounded.

alloy is the one that mattered. The chart reads `alloy.resources`; the file had
`controller.resources`, so the DaemonSet that tails every pod log on the node
shipped with nothing at all. `kubeStateMetrics` is the same trap in a different
shape -- that is the condition key, the values live under `kube-state-metrics` --
and `configReloader` in the alloy chart sits at the top level rather than under
`alloy`. Each one is verified by rendering the chart and reading the resources
back off the containers, because a values key that is ignored looks exactly
like one that works.

reloader turned out to be set and still wrong: 64Mi request against a measured
p95 of 73M, so the pod ran permanently above its own request and stayed a
standing eviction candidate. That is the pod that restarts every other pod, so
it is the last one that should be evicted. Raised to 96Mi.

Requests are set at p95 throughout, grafana, playwright and alloy included.
Left at the values first proposed they would have sat below their own p95 and
queued for eviction ahead of everything smaller. CPU limits are deliberately
absent: the node is I/O bound at 5% CPU, and CFS throttling would turn disk
wait into runnable-throttled, which is the failure mode that took the node down.

The prometheus and alertmanager configReloader sidecars are left open: chart
86.2.3 does not template the key, so reaching those two containers needs a
postRenderer.

Verified: all four charts render with the resources landing on the intended
containers, and 16/16 local gates pass.
2026-09-28 10:14:52 +02:00
forust 9b91b5847e fix(compose): stop pointing at the tag the build dropped
These five stacks still asked for :latest, but the build stopped pushing
it - on main it only pushes main and prod, on dev only dev. Every one of
these images therefore resolved only because the registry still had a
stale :latest from before that change, and the next time one of them was
actually built the reference would have dangled.

Named for dtek-notif: of the seven, only that one still resolved at
:prod, because it is the only image not rebuilt since the build dropped
:latest - and it is the only one of these five stacks the deploy does not
manage (no `active` marker, and its file is docker-compose.yaml, which
the COMPOSE_STACKS glob does not even match). Touching
dtek_notif/docker-compose.yaml matches dtek_notif/* in the build's
changed-service detector, so pushing this builds it and publishes
:prod for it too.

Verified the other six resolve at :prod in the registry.
2026-09-27 18:11:45 +02:00
forustandClaude Opus 4.8 0691536f28 fix(deploy): roll out our images by digest instead of a moving tag
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in 1505b63 therefore could not recover from a bad image.

Pin the digest at apply time. A digest is not knowable when a manifest is
written, so render_pinned resolves it on the way into the cluster and the
digest is never committed. Git keeps a readable `:prod`, Renovate keeps
seeing exactly the manifests it saw before, and the previous revision of
each workload now holds the digest that was actually serving, so undo
restores those exact bytes.

imagePullPolicy is dropped from the manifests rather than set to
IfNotPresent: a reference that is not `:latest` already defaults to it, and
that is what the Kubernetes docs ask for alongside a digest.

An unresolvable image is fatal instead of a warning, because carrying on
would quietly apply a mutable tag again.

restart_stale_images keeps its comparison but is no longer how a rebuild
reaches the cluster -- the pinned template rolls out on its own now. What
is left is a drift check for hand-run `kubectl set image`, so it matches
the container by repository: a pod's status now reports `repo@sha256:...`
while the manifest still says `:prod`.

The build job stops pushing `:latest` altogether, which removes the tag
that a dev branch could otherwise move under a prod deploy.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:09:06 +02:00
forust 2b9e34ba4a fix(k8s): add readiness probes so a bad image cannot look healthy
Six of eight workloads had no readinessProbe, so a pod turned Ready the
moment its process started. The verify job relies on `rollout status`, so
it passed for images that crash-looped or served errors, which left the
rollback safety net inert.

Each probe targets the path the service is actually reached on:
- homepages: / (verified 200)
- error-pages: /404.html, the path Traefik's errorPages middleware
  requests. / returns 403 by design and would never pass.
- webinar-checker: /health (verified 200). /metrics also answers, but it
  is a Prometheus endpoint, not a readiness signal.

The two userbot deployments stay without probes: they expose no port and
no session file, and the panel reaches Telegram through its own client. A
truthful signal there needs a health endpoint in the app itself.
2026-09-27 10:01:09 +02:00
forust 1d9a85bef9 fix(edu-master): stop 20722d false alert on zeroed last_success gauge
ci / validate (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m52s
deploy / apply-k8s (push) Successful in 1m48s
deploy / apply-compose (push) Successful in 13s
checker.py initialises last_success to 0, so right after a pod restart
`time() - last_success` equals the current epoch. The rule compared that
against 300, went firing instantly, and humanizeDuration rendered the raw
epoch delta as ~20722d. The last_run > 0 guard did not help because a run
happens long before the first success.

Guard the duration rule on last_success > 0, keeping the duration
expression on the left of `and` so $value stays the real gap, and add a
separate WebinarCheckerNeverSucceeded rule for the zeroed-gauge case so a
checker that has never succeeded is still caught.
2026-09-26 17:50:45 +02:00
renovate-bot 987a89f022 chore(deps): update container patch updates
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Canceled after 0s
ci / lint-yaml (pull_request) Canceled after 0s
ci / lint-dockerfiles (pull_request) Canceled after 0s
ci / validate (pull_request) Canceled after 0s
ci / lint-ruff (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
renovate-ci / validate-renovate (pull_request) Successful in 18s
2026-09-25 22:18:07 +00:00
forust 5cf0c90258 style: prettier formatting for edu_master alerts and postgres readme
renovate-ci / validate-renovate (push) Skipped
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 4s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 0s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 7s
2026-09-23 16:10:33 +02:00
forust 46c7e99b1d chore(deploy): rework k8s pipeline, monitoring and postgres 17
Deploy workflow uses git-tracked manifests, DISABLED flag and kustomize overlays; add webinar-checker metrics with ServiceMonitor and alerts; upgrade shared postgres to 17 with statuspage DB and probes/resources.
2026-09-23 15:47:26 +02:00
renovate-bot e502f46f43 chore(deps): update redis docker tag to v8.10.1 2026-09-14 09:46:49 +00:00
renovate-bot f5b5ecaafa chore(deps): update mcr.microsoft.com/playwright docker tag to v1.63.0
ci / lint-prettier (pull_request) Successful in 8s
ci / lint-ruff (pull_request) Successful in 4s
ci / lint-yaml (pull_request) Successful in 6s
ci / lint-dockerfiles (pull_request) Successful in 4s
ci / validate (pull_request) Successful in 5s
ci / lint-prettier (push) Successful in 7s
ci / lint-ruff (push) Successful in 3s
ci / lint-yaml (push) Successful in 6s
ci / lint-dockerfiles (push) Successful in 4s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (pull_request) Successful in 8s
ci / build (pull_request) Has been skipped
ci / build (push) Has been skipped
ci / deploy-userbot-panel (pull_request) Has been skipped
ci / deploy-userbot-panel (push) Has been skipped
2026-09-14 09:01:58 +00:00
forustandCopilot 23ed72826a chore(images): pin service image updates
Replace floating service images with reviewable tags or digests.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-14 09:16:01 +02:00
forust 6715f9e9af chore(k8s): raise resource limits for glance, edu and crowdsec 2026-09-14 01:29:28 +02:00
forustandCopilot 726b3ee544 fix(edu): run redis as statefulset
ci / deploy-userbot-panel (push) Has been skipped
ci / lint-ruff (push) Successful in 5s
ci / lint-yaml (push) Successful in 7s
ci / build (push) Successful in 3s
ci / lint-prettier (push) Successful in 11s
ci / lint-dockerfiles (push) Successful in 4s
ci / validate (push) Successful in 5s
Keep the existing Redis PVC and data while migrating the edu-master workload from Deployment to StatefulSet.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-10 23:35:59 +02:00
forust 85f05c26cb feat(edu): activate k8s management for edu-master 2026-09-10 00:55:22 +02:00
forust 68630eb773 fix(edu): read diary event times directly from DOM, drop 25 AJAX clicks
ci / lint-prettier (push) Successful in 9s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 10s
ci / lint-dockerfiles (push) Successful in 7s
ci / validate (push) Successful in 6s
ci / build (push) Has been skipped
ci / deploy-userbot-panel (push) Has been skipped
_collect_event_times re-clicked every a.event-link and waited ~3s per
event for a visible span.data, but the calendar embeds all times in
div.event-full-info[data-event-full-info-id] span.data already. The old
loop took ~109s for 25 events and collected 0 (original divs stay
sf-hidden), effectively hanging /diary. Now a single evaluate reads all
times (~3.7s), parsing HH:MM from p.date span.data.
2026-09-10 00:31:04 +02:00
forust dad9cf2104 feat(edu): parse event times in /diary and drop weekend days
ci / lint-prettier (push) Successful in 9s
ci / lint-ruff (push) Successful in 4s
ci / lint-yaml (push) Successful in 6s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 6s
ci / build (push) Has been skipped
ci / deploy-userbot-panel (push) Has been skipped
- fetch each event's time via Playwright (click event-link, read span.data, close fancybox) and render as 'title (HH:MM)'
- '08:00' placeholder renders as localized 'unknown' (time_unknown key in ru/uk/en)
- diary week view shows Mon-Fri only (title ends at Friday)
- diary month view skips Sat/Sun by weekday_idx with name-based fallback
- schedule keyboard drops Sat/Sun day buttons
2026-09-10 00:05:56 +02:00
forust 60886e8be2 style(edu): ruff-format webinar-checker/checker.py
ci / lint-prettier (push) Successful in 13s
ci / lint-ruff (push) Successful in 4s
ci / lint-yaml (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 6s
ci / build (push) Has been skipped
ci / deploy-userbot-panel (push) Has been skipped
2026-09-09 23:47:04 +02:00
forust 4528321225 fix(edu): add TZ to .env.example 2026-09-09 11:46:35 +02:00
forust b246a3dea1 fix(edu): review findings for checker.py i18n — group chat language via resolve_lang/chat: keys (default uk), weekday normalization with _norm_day + logging, drop dead translation keys, html.escape, single lang lookup, distinct whitelist emoji 2026-09-09 11:44:17 +02:00
forustandassistant 4c59a2d2fe lang(edu): translate k8s manifest comments to English only
- restore-seed-job.yaml.example: translate runbook to English
- secrets.yaml.example: translate section headers to English
- webinar-checker.yaml: translate initContainer dependency-order comments to English

Co-authored-by: assistant
2026-09-09 10:46:36 +02:00
forust 30e6f05584 fix(edu_master): satisfy ruff in schedule scraper
- rename ambiguous loop var, merge nested if (E741, SIM102)
- use tempfile.gettempdir() for debug dump (S108)
- drop unused total_lessons assignment (F841)
2026-09-06 20:37:18 +02:00
forust bc8d74fd28 feat(schedule): per-user class picker, parser fixes
ci / lint-prettier (push) Successful in 8s
ci / lint-ruff (push) Failing after 5s
ci / lint-yaml (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 4s
ci / validate (push) Successful in 4s
ci / build (push) Has been skipped
- fix parser: capture tr attrs via finditer, strip HTML comments before
  cell parse (was leaving '-->' in subject names)
- store class choice in redis: user:{id}:schedule_class (private) and
  chat:{id}:schedule_class (groups, admin-only via /setclass)
- /schedule renders day for stored class, /setclass sets it directly
- drop teacher emoji, format grade as "N клас"
2026-09-06 16:17:48 +02:00
forust d2d4efb0d7 feat: add schedule scraper for lessons table
ci / lint-prettier (push) Successful in 9s
ci / lint-ruff (push) Failing after 4s
ci / lint-yaml (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 7s
ci / build (push) Has been skipped
Add /schedule command to scrape edu.edu.vn.ua/lessons/table via Playwright.
Inline keyboard flow: pick weekday (with today/tomorrow shortcuts),
then pick class. Cache 5h per user. Parse subjects/notes/teachers,
multi-lesson cells (hr-separated).
2026-09-06 15:46:41 +02:00
forust b752bf88bd feat: add edu_master k8s manifests for k0s migration
ci / lint-prettier (push) Successful in 10s
ci / lint-ruff (push) Successful in 4s
ci / lint-yaml (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 5s
ci / build (push) Successful in 35s
2026-09-06 15:00:51 +02:00
forust a2df8504f5 Fix DL3013: pin pip version in webinar-checker Dockerfile
lint / prettier (push) Successful in 8s
lint / ruff (push) Successful in 4s
lint / yamllint (push) Successful in 6s
lint / hadolint (push) Successful in 4s
validate / yaml (push) Successful in 6s
validate / k8s (push) Successful in 4s
2026-06-21 22:07:44 +02:00
forust fb43306571 Fix all lint issues: Dockerfiles (DL3015/DL3013/DL4006) + Ruff (173→0 errors)
lint / prettier (push) Successful in 8s
lint / ruff (push) Successful in 5s
lint / yamllint (push) Successful in 7s
lint / hadolint (push) Failing after 4s
validate / yaml (push) Successful in 5s
validate / k8s (push) Successful in 5s
Dockerfile fixes:
- edu_master/phpsessid-bot: add --no-install-recommends, pin pip versions
- edu_master/webinar-checker: pin pip versions with --no-cache-dir
- userbot: add SHELL with pipefail for pipe operations

Ruff fixes (173 → 0):
- W293/W291/W292: whitespace clean via ruff format
- N806: camelCase → snake_case (anilist, safone, hearts, flux, etc.)
- ARG001/ARG002: prefix unused params with _
- SIM115: use context managers for file I/O
- SIM117: combine nested with statements
- S608: noqa on SQL f-strings (module name is validated)
- E402/N812/N817: import fixes
- B023: pass loop variable as argument
- I001: auto-sorted imports
- syntax: fixed = vs == in dtek_notif/main.py
2026-06-21 22:01:51 +02:00
forust 227e5fda27 chore: batch lint fixes across userbot and edu_master
- S113: Add timeout=10 to all requests calls (74 fixes)
- E722: Replace bare except: with except Exception:
- B904: Replace redundant re-raise with bare raise
- E402: Add noqa for intentional late imports after import_library()
- S102/S307/S310/S311/S603/S605/S606/S607/S108: Add noqa for intentional usage
- F601: Fix duplicate dict key in unsplash.py
- N802: Rename ReplyCheck -> reply_check with backward compat alias
- N813: Rename bs -> BS in icons.py
- B007/B020: Rename loop var _j in animations.py
- SIM102: Collapse nested if in autofwd.py
- SIM113: Use enumerate() in calculator.py
- A002: Add noqa for builtin shadowing in admlist.py
- F811: Add noqa for cohere redefinition
- edu_master: Fix ARG001, S108, S110, SIM117, apply --unsafe-fixes
- Add modules_list.txt with full module inventory
2026-06-19 15:36:25 +02:00
forust 2de6131ba7 chore: apply prettier formatting across compose files and configs 2026-06-19 12:03:11 +02:00
forust a68c01a68f chore: add gitea container registry compose support for localy builded apps
Deploy to Server / deploy (push) Has been cancelled
2026-06-07 14:37:57 +02:00
forust ea286e117d feat: add diary (/diary) command with calendar parsing via Playwright
Deploy to Server / deploy (push) Has been cancelled
- Parse school diary calendar HTML table for daily/weekly/monthly views
- Cache diary data with 5-minute TTL
- Inline keyboard for today/tomorrow/week/month selection
- Ukrainian month/weekday names and formatting
2026-05-27 19:39:45 +02:00
forust d7c05fd058 fix: use webinar links to detect duplicates, supress logspam from telegram bot 2026-01-19 17:01:23 +01:00
forust 81592b6142 feat: add group support for webinar notifier 2026-01-13 00:24:04 +01:00
forust 9480576966 fix: update bots' code to match .env keys 2026-01-12 15:33:42 +01:00
forust 6cb49be951 chore: update gitignore, add translations 2025-11-27 18:58:17 +01:00
forust 3bcabcce4b feat: notify on new webinar, storing N last webinars in redis 2025-11-27 16:33:15 +01:00
forust a54b7b000f fix: add timeout for playwright to load page and debug response saving 2025-11-27 15:50:12 +01:00
forust 380288d103 fix: port conflicts 2025-11-26 22:55:16 +01:00
forust 8fae8bf75d feat: add admin commands to manage whitelist, extended webinar reporting 2025-11-26 22:25:47 +01:00
forust d4e0e7f37a feat: telegram webianr checker bot 2025-11-26 22:12:03 +01:00
forust cb92bff0c5 feat: basic webinar checker bot 2025-11-26 21:08:43 +01:00
forust bfbe841154 feat: redis integration 2025-11-26 20:53:53 +01:00
forust 7cbc5bf4f3 feat: edu phpsessid updater bot 2025-11-26 17:00:44 +01:00
forust fd3d2affe7 full refactor
edu_master
glance
n8n
nextcloud
portainer
traefik

gitignore comments
2025-11-15 13:30:51 +01:00
forust 75874a2ff3 Refactor Docker Compose files for n8n and Portainer: update volume paths and network configurations; remove deprecated n8n configuration file. 2025-11-13 03:38:48 +01:00
forust 0403524992 Remove deprecated Docker Compose files and configurations for various services including edu_master, glance, homer, and traefik. Introduce new Docker Compose files for dockmon, glance, homer, and traefik with updated configurations and environment variables. Ensure all services are integrated with Traefik for routing and TLS management. 2025-11-11 01:32:14 +01:00
forust 32aac89fdf init, .gitignore 2025-11-11 00:02:49 +01:00