fix(crowdsec): stop the bouncer from 403-ing our own deploys
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 32s
ci / build (push) Successful in 1s
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 32s
ci / build (push) Successful in 1s
Every bouncer-protected request blocks on a synchronous GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403 when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi on a single replica, so idle lookups measured 1.3-7.4s and a deploy burst pushed them past the fork's implicit 10s default: gitea answered 403 for ten seconds straight, and containerd turned those 403s on gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled pods. Three changes, plus the 403 feedback loop that made it sticky: * gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route. A deploy fires hundreds of parallel authenticated OCI requests (runner Action API, manifest inspect per own image, containerd pulls, smoke probes) and scanners gain nothing from a registry that already does its own token auth. The web UI route keeps the bouncer. * crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs), so a second replica would corrupt the decision store. * crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang. LePresidente/http-generic-403-bf then banned us for our own 403s: five POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT address 192.168.88.1 that the Gitea Actions runner presents to Traefik is not covered by the home-dynamic-IP whitelist. That scenario cannot be dropped per-scenario - it is baked into the hub item crowdsecurity/http-generic-bf v0.9, and disabling the whole base-http-scenarios collection would cost ~40 useful detections. So the janitor now deletes its decisions hourly and a new forust/lan whitelist postoverflow covers 192.168.88.0/24. First janitor run removed 165 decisions; none of the remaining ones are local. Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth challenge, LAPI at 60m CPU with no throttling.
This commit is contained in:
1 parent
6a9a460769
commit
f7cd75d65e
4 files changed
+68
-8
No files matched your search
@@ -68,6 +68,23 @@ config:
|
||||
reason: "Home dynamic IP"
|
||||
expression:
|
||||
- evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz")
|
||||
# The hairpin-NAT address of the router (192.168.88.1) is what the
|
||||
# Gitea Actions runner presents to Traefik - it is NOT the home
|
||||
# dynamic IP, so the whitelist above did not cover it. During a
|
||||
# deploy the runner POSTs to the Actions API many times a second;
|
||||
# a single 403 storm was enough to earn it a 4h ban and break every
|
||||
# later job. Whitelisting the whole LAN also covers phones and
|
||||
# tablets browsing over 192.168.88.0/24.
|
||||
lan.yaml: |
|
||||
name: forust/lan
|
||||
description: "Whitelist local network"
|
||||
whitelist:
|
||||
reason: "Local network"
|
||||
cidr:
|
||||
- "127.0.0.0/8"
|
||||
- "10.0.0.0/8"
|
||||
- "172.16.0.0/12"
|
||||
- "192.168.0.0/16"
|
||||
|
||||
lapi:
|
||||
env:
|
||||
@@ -90,13 +107,20 @@ lapi:
|
||||
enabled: true
|
||||
size: 1Gi
|
||||
storageClassName: local-path-retain
|
||||
# LAPI answers a blocking /v1/decisions lookup for EVERY bouncer-protected
|
||||
# request (whole Traefik front door), so it is the hot path of the proxy.
|
||||
# At 400m/500Mi it went CPU-throttled and idle lookups measured 1.3-7.4s,
|
||||
# which pushed requests into the bouncer's fail-closed 403.
|
||||
# Single replica on purpose: LAPI is stateful (BoltDB on the `data` PVC,
|
||||
# credentials on the `config` PVC) - two replicas sharing those RWO
|
||||
# volumes would corrupt the decision store. Scale up CPU, not replicas.
|
||||
resources:
|
||||
limits:
|
||||
cpu: 400m
|
||||
memory: 500Mi
|
||||
cpu: 1500m
|
||||
memory: 1Gi
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 150Mi
|
||||
cpu: 250m
|
||||
memory: 500Mi
|
||||
service:
|
||||
type: ClusterIP
|
||||
storeLAPICscliCredentialsInSecret: true
|
||||
Reference in new issue
Block a user