Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.
Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.
Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.
prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.
Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.
CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.
Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
Eighteen containers had no memory limit at all, so nothing on the node could
bound them. Three of the values files even claimed to set resources: Helm does
not complain about a key it does not recognise, so the block sat there looking
like a limit while the pod ran unbounded.
alloy is the one that mattered. The chart reads `alloy.resources`; the file had
`controller.resources`, so the DaemonSet that tails every pod log on the node
shipped with nothing at all. `kubeStateMetrics` is the same trap in a different
shape -- that is the condition key, the values live under `kube-state-metrics` --
and `configReloader` in the alloy chart sits at the top level rather than under
`alloy`. Each one is verified by rendering the chart and reading the resources
back off the containers, because a values key that is ignored looks exactly
like one that works.
reloader turned out to be set and still wrong: 64Mi request against a measured
p95 of 73M, so the pod ran permanently above its own request and stayed a
standing eviction candidate. That is the pod that restarts every other pod, so
it is the last one that should be evicted. Raised to 96Mi.
Requests are set at p95 throughout, grafana, playwright and alloy included.
Left at the values first proposed they would have sat below their own p95 and
queued for eviction ahead of everything smaller. CPU limits are deliberately
absent: the node is I/O bound at 5% CPU, and CFS throttling would turn disk
wait into runnable-throttled, which is the failure mode that took the node down.
The prometheus and alertmanager configReloader sidecars are left open: chart
86.2.3 does not template the key, so reaching those two containers needs a
postRenderer.
Verified: all four charts render with the resources landing on the intended
containers, and 16/16 local gates pass.
These five stacks still asked for :latest, but the build stopped pushing
it - on main it only pushes main and prod, on dev only dev. Every one of
these images therefore resolved only because the registry still had a
stale :latest from before that change, and the next time one of them was
actually built the reference would have dangled.
Named for dtek-notif: of the seven, only that one still resolved at
:prod, because it is the only image not rebuilt since the build dropped
:latest - and it is the only one of these five stacks the deploy does not
manage (no `active` marker, and its file is docker-compose.yaml, which
the COMPOSE_STACKS glob does not even match). Touching
dtek_notif/docker-compose.yaml matches dtek_notif/* in the build's
changed-service detector, so pushing this builds it and publishes
:prod for it too.
Verified the other six resolve at :prod in the registry.
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in 1505b63 therefore could not recover from a bad image.
Pin the digest at apply time. A digest is not knowable when a manifest is
written, so render_pinned resolves it on the way into the cluster and the
digest is never committed. Git keeps a readable `:prod`, Renovate keeps
seeing exactly the manifests it saw before, and the previous revision of
each workload now holds the digest that was actually serving, so undo
restores those exact bytes.
imagePullPolicy is dropped from the manifests rather than set to
IfNotPresent: a reference that is not `:latest` already defaults to it, and
that is what the Kubernetes docs ask for alongside a digest.
An unresolvable image is fatal instead of a warning, because carrying on
would quietly apply a mutable tag again.
restart_stale_images keeps its comparison but is no longer how a rebuild
reaches the cluster -- the pinned template rolls out on its own now. What
is left is a drift check for hand-run `kubectl set image`, so it matches
the container by repository: a pod's status now reports `repo@sha256:...`
while the manifest still says `:prod`.
The build job stops pushing `:latest` altogether, which removes the tag
that a dev branch could otherwise move under a prod deploy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Six of eight workloads had no readinessProbe, so a pod turned Ready the
moment its process started. The verify job relies on `rollout status`, so
it passed for images that crash-looped or served errors, which left the
rollback safety net inert.
Each probe targets the path the service is actually reached on:
- homepages: / (verified 200)
- error-pages: /404.html, the path Traefik's errorPages middleware
requests. / returns 403 by design and would never pass.
- webinar-checker: /health (verified 200). /metrics also answers, but it
is a Prometheus endpoint, not a readiness signal.
The two userbot deployments stay without probes: they expose no port and
no session file, and the panel reaches Telegram through its own client. A
truthful signal there needs a health endpoint in the app itself.
checker.py initialises last_success to 0, so right after a pod restart
`time() - last_success` equals the current epoch. The rule compared that
against 300, went firing instantly, and humanizeDuration rendered the raw
epoch delta as ~20722d. The last_run > 0 guard did not help because a run
happens long before the first success.
Guard the duration rule on last_success > 0, keeping the duration
expression on the left of `and` so $value stays the real gap, and add a
separate WebinarCheckerNeverSucceeded rule for the zeroed-gauge case so a
checker that has never succeeded is still caught.
Deploy workflow uses git-tracked manifests, DISABLED flag and kustomize overlays; add webinar-checker metrics with ServiceMonitor and alerts; upgrade shared postgres to 17 with statuspage DB and probes/resources.
Keep the existing Redis PVC and data while migrating the edu-master workload from Deployment to StatefulSet.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
_collect_event_times re-clicked every a.event-link and waited ~3s per
event for a visible span.data, but the calendar embeds all times in
div.event-full-info[data-event-full-info-id] span.data already. The old
loop took ~109s for 25 events and collected 0 (original divs stay
sf-hidden), effectively hanging /diary. Now a single evaluate reads all
times (~3.7s), parsing HH:MM from p.date span.data.
- fetch each event's time via Playwright (click event-link, read span.data, close fancybox) and render as 'title (HH:MM)'
- '08:00' placeholder renders as localized 'unknown' (time_unknown key in ru/uk/en)
- diary week view shows Mon-Fri only (title ends at Friday)
- diary month view skips Sat/Sun by weekday_idx with name-based fallback
- schedule keyboard drops Sat/Sun day buttons
- restore-seed-job.yaml.example: translate runbook to English
- secrets.yaml.example: translate section headers to English
- webinar-checker.yaml: translate initContainer dependency-order comments to English
Co-authored-by: assistant
- rename ambiguous loop var, merge nested if (E741, SIM102)
- use tempfile.gettempdir() for debug dump (S108)
- drop unused total_lessons assignment (F841)
- fix parser: capture tr attrs via finditer, strip HTML comments before
cell parse (was leaving '-->' in subject names)
- store class choice in redis: user:{id}:schedule_class (private) and
chat:{id}:schedule_class (groups, admin-only via /setclass)
- /schedule renders day for stored class, /setclass sets it directly
- drop teacher emoji, format grade as "N клас"
- S113: Add timeout=10 to all requests calls (74 fixes)
- E722: Replace bare except: with except Exception:
- B904: Replace redundant re-raise with bare raise
- E402: Add noqa for intentional late imports after import_library()
- S102/S307/S310/S311/S603/S605/S606/S607/S108: Add noqa for intentional usage
- F601: Fix duplicate dict key in unsplash.py
- N802: Rename ReplyCheck -> reply_check with backward compat alias
- N813: Rename bs -> BS in icons.py
- B007/B020: Rename loop var _j in animations.py
- SIM102: Collapse nested if in autofwd.py
- SIM113: Use enumerate() in calculator.py
- A002: Add noqa for builtin shadowing in admlist.py
- F811: Add noqa for cohere redefinition
- edu_master: Fix ARG001, S108, S110, SIM117, apply --unsafe-fixes
- Add modules_list.txt with full module inventory
- Parse school diary calendar HTML table for daily/weekly/monthly views
- Cache diary data with 5-minute TTL
- Inline keyboard for today/tomorrow/week/month selection
- Ukrainian month/weekday names and formatting