Files
homelab/crowdsec/k8s/crowdsec-values.yaml
T
forust f7cd75d65e
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 32s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / build (push) Successful in 1s
fix(crowdsec): stop the bouncer from 403-ing our own deploys
Every bouncer-protected request blocks on a synchronous
GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403
when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi
on a single replica, so idle lookups measured 1.3-7.4s and a deploy
burst pushed them past the fork's implicit 10s default: gitea answered
403 for ten seconds straight, and containerd turned those 403s on
gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled
pods.

Three changes, plus the 403 feedback loop that made it sticky:

* gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route.
  A deploy fires hundreds of parallel authenticated OCI requests
  (runner Action API, manifest inspect per own image, containerd pulls,
  smoke probes) and scanners gain nothing from a registry that already
  does its own token auth. The web UI route keeps the bouncer.
* crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on
  purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs),
  so a second replica would corrupt the decision store.
* crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the
  implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang.

LePresidente/http-generic-403-bf then banned us for our own 403s: five
POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT
address 192.168.88.1 that the Gitea Actions runner presents to Traefik
is not covered by the home-dynamic-IP whitelist. That scenario cannot
be dropped per-scenario - it is baked into the hub item
crowdsecurity/http-generic-bf v0.9, and disabling the whole
base-http-scenarios collection would cost ~40 useful detections. So the
janitor now deletes its decisions hourly and a new forust/lan
whitelist postoverflow covers 192.168.88.0/24. First janitor run
removed 165 decisions; none of the remaining ones are local.

Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way
parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth
challenge, LAPI at 60m CPU with no throttling.
2026-09-27 17:08:22 +02:00

127 lines
3.9 KiB
YAML

container_runtime: containerd
agent:
acquisition: []
additionalAcquisition:
- labels:
type: traefik
limit: 1000
query: |
{namespace="traefik"}
source: loki
url: http://loki.prometheus.svc.cluster.local:3100/
wait_for_ready: 30s
env:
- name: COLLECTIONS
value: crowdsecurity/traefik crowdsecurity/base-http-scenarios
- name: DISABLE_COLLECTIONS
value: crowdsecurity/sshd
metrics:
enabled: true
serviceMonitor:
additionalLabels:
release: prometheus-stack
enabled: true
# Static machine identity: agent pods mount pre-created LAPI credentials
# (Secret crowdsec-agent-credentials, key local_api_credentials.yaml)
# at the exact path the agent entrypoint expects. Together with the
# patched register-init (enforced by janitor-cronjob.yaml) the agent
# never calls `cscli lapi register` in steady state, so pod names,
# restarts and reboots can no longer break it.
extraVolumes:
- name: static-creds
secret:
secretName: crowdsec-agent-credentials
items:
- key: local_api_credentials.yaml
path: local_api_credentials.yaml
extraVolumeMounts:
- name: static-creds
mountPath: /tmp_config/local_api_credentials.yaml
subPath: local_api_credentials.yaml
readOnly: true
resources:
limits:
cpu: 200m
memory: 500Mi
requests:
cpu: 50m
memory: 100Mi
config:
parsers:
s02-enrich:
mobile-whitelist.yaml: |
name: forust/mobile-whitelist
description: "Whitelist SWAN/4ka mobile network"
whitelist:
reason: "Mobile IP whitelist"
cidr:
- "84.245.64.0/18"
postoverflows:
s01-whitelist:
home-dynamic-ip.yaml: |
name: forust/home-dynamic-ip
description: "Whitelist home dynamic IP"
whitelist:
reason: "Home dynamic IP"
expression:
- evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz")
# The hairpin-NAT address of the router (192.168.88.1) is what the
# Gitea Actions runner presents to Traefik - it is NOT the home
# dynamic IP, so the whitelist above did not cover it. During a
# deploy the runner POSTs to the Actions API many times a second;
# a single 403 storm was enough to earn it a 4h ban and break every
# later job. Whitelisting the whole LAN also covers phones and
# tablets browsing over 192.168.88.0/24.
lan.yaml: |
name: forust/lan
description: "Whitelist local network"
whitelist:
reason: "Local network"
cidr:
- "127.0.0.0/8"
- "10.0.0.0/8"
- "172.16.0.0/12"
- "192.168.0.0/16"
lapi:
env:
- name: COLLECTIONS
value: crowdsecurity/traefik crowdsecurity/base-http-scenarios
- name: DISABLE_COLLECTIONS
value: crowdsecurity/linux crowdsecurity/sshd
metrics:
enabled: true
serviceMonitor:
additionalLabels:
release: prometheus-stack
enabled: true
persistentVolume:
config:
enabled: true
size: 100Mi
storageClassName: local-path-retain
data:
enabled: true
size: 1Gi
storageClassName: local-path-retain
# LAPI answers a blocking /v1/decisions lookup for EVERY bouncer-protected
# request (whole Traefik front door), so it is the hot path of the proxy.
# At 400m/500Mi it went CPU-throttled and idle lookups measured 1.3-7.4s,
# which pushed requests into the bouncer's fail-closed 403.
# Single replica on purpose: LAPI is stateful (BoltDB on the `data` PVC,
# credentials on the `config` PVC) - two replicas sharing those RWO
# volumes would corrupt the decision store. Scale up CPU, not replicas.
resources:
limits:
cpu: 1500m
memory: 1Gi
requests:
cpu: 250m
memory: 500Mi
service:
type: ClusterIP
storeLAPICscliCredentialsInSecret: true