ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 32s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / build (push) Successful in 1s
Every bouncer-protected request blocks on a synchronous GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403 when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi on a single replica, so idle lookups measured 1.3-7.4s and a deploy burst pushed them past the fork's implicit 10s default: gitea answered 403 for ten seconds straight, and containerd turned those 403s on gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled pods. Three changes, plus the 403 feedback loop that made it sticky: * gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route. A deploy fires hundreds of parallel authenticated OCI requests (runner Action API, manifest inspect per own image, containerd pulls, smoke probes) and scanners gain nothing from a registry that already does its own token auth. The web UI route keeps the bouncer. * crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs), so a second replica would corrupt the decision store. * crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang. LePresidente/http-generic-403-bf then banned us for our own 403s: five POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT address 192.168.88.1 that the Gitea Actions runner presents to Traefik is not covered by the home-dynamic-IP whitelist. That scenario cannot be dropped per-scenario - it is baked into the hub item crowdsecurity/http-generic-bf v0.9, and disabling the whole base-http-scenarios collection would cost ~40 useful detections. So the janitor now deletes its decisions hourly and a new forust/lan whitelist postoverflow covers 192.168.88.0/24. First janitor run removed 165 decisions; none of the remaining ones are local. Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth challenge, LAPI at 60m CPU with no throttling.
127 lines
3.9 KiB
YAML
127 lines
3.9 KiB
YAML
container_runtime: containerd
|
|
|
|
agent:
|
|
acquisition: []
|
|
additionalAcquisition:
|
|
- labels:
|
|
type: traefik
|
|
limit: 1000
|
|
query: |
|
|
{namespace="traefik"}
|
|
source: loki
|
|
url: http://loki.prometheus.svc.cluster.local:3100/
|
|
wait_for_ready: 30s
|
|
env:
|
|
- name: COLLECTIONS
|
|
value: crowdsecurity/traefik crowdsecurity/base-http-scenarios
|
|
- name: DISABLE_COLLECTIONS
|
|
value: crowdsecurity/sshd
|
|
metrics:
|
|
enabled: true
|
|
serviceMonitor:
|
|
additionalLabels:
|
|
release: prometheus-stack
|
|
enabled: true
|
|
# Static machine identity: agent pods mount pre-created LAPI credentials
|
|
# (Secret crowdsec-agent-credentials, key local_api_credentials.yaml)
|
|
# at the exact path the agent entrypoint expects. Together with the
|
|
# patched register-init (enforced by janitor-cronjob.yaml) the agent
|
|
# never calls `cscli lapi register` in steady state, so pod names,
|
|
# restarts and reboots can no longer break it.
|
|
extraVolumes:
|
|
- name: static-creds
|
|
secret:
|
|
secretName: crowdsec-agent-credentials
|
|
items:
|
|
- key: local_api_credentials.yaml
|
|
path: local_api_credentials.yaml
|
|
extraVolumeMounts:
|
|
- name: static-creds
|
|
mountPath: /tmp_config/local_api_credentials.yaml
|
|
subPath: local_api_credentials.yaml
|
|
readOnly: true
|
|
resources:
|
|
limits:
|
|
cpu: 200m
|
|
memory: 500Mi
|
|
requests:
|
|
cpu: 50m
|
|
memory: 100Mi
|
|
|
|
config:
|
|
parsers:
|
|
s02-enrich:
|
|
mobile-whitelist.yaml: |
|
|
name: forust/mobile-whitelist
|
|
description: "Whitelist SWAN/4ka mobile network"
|
|
whitelist:
|
|
reason: "Mobile IP whitelist"
|
|
cidr:
|
|
- "84.245.64.0/18"
|
|
|
|
postoverflows:
|
|
s01-whitelist:
|
|
home-dynamic-ip.yaml: |
|
|
name: forust/home-dynamic-ip
|
|
description: "Whitelist home dynamic IP"
|
|
whitelist:
|
|
reason: "Home dynamic IP"
|
|
expression:
|
|
- evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz")
|
|
# The hairpin-NAT address of the router (192.168.88.1) is what the
|
|
# Gitea Actions runner presents to Traefik - it is NOT the home
|
|
# dynamic IP, so the whitelist above did not cover it. During a
|
|
# deploy the runner POSTs to the Actions API many times a second;
|
|
# a single 403 storm was enough to earn it a 4h ban and break every
|
|
# later job. Whitelisting the whole LAN also covers phones and
|
|
# tablets browsing over 192.168.88.0/24.
|
|
lan.yaml: |
|
|
name: forust/lan
|
|
description: "Whitelist local network"
|
|
whitelist:
|
|
reason: "Local network"
|
|
cidr:
|
|
- "127.0.0.0/8"
|
|
- "10.0.0.0/8"
|
|
- "172.16.0.0/12"
|
|
- "192.168.0.0/16"
|
|
|
|
lapi:
|
|
env:
|
|
- name: COLLECTIONS
|
|
value: crowdsecurity/traefik crowdsecurity/base-http-scenarios
|
|
- name: DISABLE_COLLECTIONS
|
|
value: crowdsecurity/linux crowdsecurity/sshd
|
|
metrics:
|
|
enabled: true
|
|
serviceMonitor:
|
|
additionalLabels:
|
|
release: prometheus-stack
|
|
enabled: true
|
|
persistentVolume:
|
|
config:
|
|
enabled: true
|
|
size: 100Mi
|
|
storageClassName: local-path-retain
|
|
data:
|
|
enabled: true
|
|
size: 1Gi
|
|
storageClassName: local-path-retain
|
|
# LAPI answers a blocking /v1/decisions lookup for EVERY bouncer-protected
|
|
# request (whole Traefik front door), so it is the hot path of the proxy.
|
|
# At 400m/500Mi it went CPU-throttled and idle lookups measured 1.3-7.4s,
|
|
# which pushed requests into the bouncer's fail-closed 403.
|
|
# Single replica on purpose: LAPI is stateful (BoltDB on the `data` PVC,
|
|
# credentials on the `config` PVC) - two replicas sharing those RWO
|
|
# volumes would corrupt the decision store. Scale up CPU, not replicas.
|
|
resources:
|
|
limits:
|
|
cpu: 1500m
|
|
memory: 1Gi
|
|
requests:
|
|
cpu: 250m
|
|
memory: 500Mi
|
|
service:
|
|
type: ClusterIP
|
|
storeLAPICscliCredentialsInSecret: true
|