fix(crowdsec): stop the 403 loop that banned our own VPS and runner

Chasing why apply-k8s kept dying mid-run turned up a self-inflicted
ban loop. 585 of 586 LePresidente/http-generic-403-bf alerts in the LAPI
came from 193.181.211.79 - our own VPS - POSTing
/management.ManagementService/GetServerKey, i.e. the NetBird client's
own management call. netbird-server had never seen a single one of them,
so the 403 was not NetBird's: the bouncer was rejecting the request
before it got there. A banned peer keeps retrying, each retry is another
403, and the scenario turns five 403s in ten seconds into a 4h ban, so
the loop kept re-arming the ban it was serving. The hourly janitor step
that deleted those decisions hourly was masking all of it.

* netbird/k8s/ingress.yaml - drop the bouncer from the mesh API routes
  (gRPC-gateway management, signal, relay, /api, /oauth2). A ban there
  locks a peer out of the network it needs to reach anything else, and
  those endpoints authenticate by NetBird token, not by a login form.
  The dashboard keeps the bouncer; it is a real login surface.
  netbird-local was already exempt, so this makes prod match.

* crowdsec-middleware.yaml - CrowdsecMode stream instead of live. live
  blocked on GET /v1/decisions per request, so a burst saturated the
  LAPI and the plugin 403'd IPs that were never banned. v1.3.3 ignores
  UpdateMaxFailure in live, so fail-open is only reachable in stream;
  -1 now means an unreachable LAPI degrades to "no protection" rather
  than "every site 403". 15s poll instead of the 60s default, because
  the runner shares one public IP with the house.

  Also corrects the key name: HTTPTimeoutSeconds, not
  CrowdsecLapiTimeout, which never existed and was being silently
  dropped, leaving the 10s default. Back at 10, not the 2s f7cd75d
  guessed - a pull that times out leaves the ban cache frozen at its
  startup contents, so new bans would silently never apply.

* crowdsec-values.yaml - CIDR allowlisting moves to parsers/s02-enrich,
  where CrowdSec's docs put it: a parser whitelist drops the event before
  it reaches a bucket, so those addresses never become a decision at
  all. The old postoverflow LAN list was checked only after the ban
  existed, which is the window the deploy kept landing in. Added
  100.64.0.0/10, which the RFC 1918 blocks miss and where the
  workstation, the k0s node and the VPS actually live. The DDNS
  home-IP whitelist stays in postoverflows, because resolving a hostname
  is the expensive check the docs reserve that stage for.

  ClientTrustedIPs mirrors that list so the bouncer skips the LAPI
  round-trip entirely for those addresses.

* janitor-cronjob.yaml - drop step 5. The bouncer can no longer
  manufacture 403s, so the only remaining firings of that scenario are
  real scanners, and deleting their decisions hourly was undoing a
  working ban.

* Also lands the LAPI config.yaml.local (SQLite WAL, Central API off,
  bounded flush) that f7cd75d's comments referenced but never included:
  the LAPI was blocked in fsync on its rollback journal, and the CAPI
  resolver held a write transaction while timing out against a host
  this network cannot reach.

Verified in-cluster: no new 403-bf alerts in the 3.5min after applying,
the management endpoint answers 404 from the backend instead of 403 from
the bouncer in 0.17s, and the VPS client reports Management and Signal
connected with 2/2 relays.
This commit is contained in:
forust committed 2026-09-27 18:09:07 +02:00
1 parent f7cd75d65e
commit 892790822d
5 files changed
+141 -66

No files matched your search

+34 -11
View File
@@ -8,17 +8,40 @@ spec:
crowdsec-bouncer: crowdsec-bouncer:
enabled: true enabled: true
LogLevel: INFO LogLevel: INFO
CrowdsecMode: live # `live` blocked on a `GET /v1/decisions` per request, so a burst
# saturated the LAPI and the plugin 403'd IPs that were never banned.
# v1.3.3 ignores UpdateMaxFailure in `live`, so fail-open is only
# reachable in stream mode, which polls into a cache instead - no
# per-request call to saturate. 15s rather than the 60s default: the
# deploy runner shares one public IP with the house, so this bounds
# both how late a ban lands and how long a lifted one lingers.
CrowdsecMode: stream
UpdateIntervalSeconds: 15
# -1 = never block because the LAPI is unreachable. In v1.3.3
# handleStreamTicker only clears isCrowdsecStreamHealthy when
# updateMaxFailure != -1, and ServeHTTP 403s once it is false, so this
# makes a CrowdSec outage mean "no protection", not "every site 403".
UpdateMaxFailure: -1
CrowdsecLapiScheme: http CrowdsecLapiScheme: http
CrowdsecLapiHost: crowdsec-service.crowdsec.svc.cluster.local:8080 CrowdsecLapiHost: crowdsec-service.crowdsec.svc.cluster.local:8080
CrowdsecLapiKeyFile: "/etc/traefik/secrets/traefik-api-key" CrowdsecLapiKeyFile: "/etc/traefik/secrets/traefik-api-key"
# LAPI lookup is SYNCHRONOUS and per-request: the plugin blocks on # Bypasses the bouncer and the decision cache, no LAPI round-trip.
# `GET /v1/decisions?ip=...&banned=true` before the request reaches # Keep in sync with forust/local-network in crowdsec-values.yaml.
# the backend, and fails CLOSED (403) if the lookup exceeds the ClientTrustedIPs:
# timeout. Unset, the fork defaults to 10s, which is an eternity for - "127.0.0.0/8"
# a request path: a single slow LAPI (idle 1.3-7.4s here) turned - "10.0.0.0/8"
# every request into a 10s hang and then a self-inflicted 403. - "172.16.0.0/12"
# 2s keeps the fail-closed path fast and bounded; with the LAPI - "192.168.0.0/16"
# resourced properly (see crowdsec-values.yaml) the lookup is - "100.64.0.0/10"
# sub-100ms and this budget is never hit. - "169.254.0.0/16"
CrowdsecLapiTimeout: "2s" - "fc00::/7"
- "fe80::/10"
# The name is HTTPTimeoutSeconds, an int in seconds (min 1) - there is
# no CrowdsecLapiTimeout, and an unrecognised key is silently dropped,
# which is how this sat at the 10s default. Nothing rides on it per
# request any more, so this only bounds the stream pull - and too low
# is the dangerous direction: the LAPI needs ~2s to answer
# /v1/decisions/stream, and a pull that times out leaves the ban cache
# frozen at its startup contents ("failed sending new decisions"),
# i.e. new bans silently never apply. Keep it above the pull latency.
HTTPTimeoutSeconds: 10
+84 -20
View File
@@ -58,26 +58,15 @@ config:
reason: "Mobile IP whitelist" reason: "Mobile IP whitelist"
cidr: cidr:
- "84.245.64.0/18" - "84.245.64.0/18"
# CrowdSec's own guidance: CIDR allowlisting belongs at the parser stage.
postoverflows: # A parser whitelist discards the event before it reaches a bucket, so
s01-whitelist: # these addresses never produce an overflow and never become a decision.
home-dynamic-ip.yaml: | # A postoverflow whitelist is checked only *after* the ban exists, and
name: forust/home-dynamic-ip # the bouncer answers 403 for as long as it does - which is a window we
description: "Whitelist home dynamic IP" # do not want the deploy sitting in.
whitelist: local-network.yaml: |
reason: "Home dynamic IP" name: forust/local-network
expression: description: "Whitelist loopback, private and VPN networks"
- evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz")
# The hairpin-NAT address of the router (192.168.88.1) is what the
# Gitea Actions runner presents to Traefik - it is NOT the home
# dynamic IP, so the whitelist above did not cover it. During a
# deploy the runner POSTs to the Actions API many times a second;
# a single 403 storm was enough to earn it a 4h ban and break every
# later job. Whitelisting the whole LAN also covers phones and
# tablets browsing over 192.168.88.0/24.
lan.yaml: |
name: forust/lan
description: "Whitelist local network"
whitelist: whitelist:
reason: "Local network" reason: "Local network"
cidr: cidr:
@@ -85,6 +74,81 @@ config:
- "10.0.0.0/8" - "10.0.0.0/8"
- "172.16.0.0/12" - "172.16.0.0/12"
- "192.168.0.0/16" - "192.168.0.0/16"
# CGNAT range (RFC 6598). The workstation and the k0s node live
# here on WireGuard, and 100.64.0.0/10 is not covered by the
# RFC 1918 blocks above.
- "100.64.0.0/10"
- "169.254.0.0/16"
- "fc00::/7"
- "fe80::/10"
postoverflows:
s01-whitelist:
# The one whitelist that has to stay here: resolving a hostname is a
# network call, and the docs put expensive lookups in postoverflows on
# purpose - it runs only when a bucket actually overflows.
# ddns.forust.xyz is the public home address, not a private one, so
# forust/local-network does not cover it.
home-dynamic-ip.yaml: |
name: forust/home-dynamic-ip
description: "Whitelist home dynamic IP"
whitelist:
reason: "Home dynamic IP"
expression:
- evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz")
# LAPI-only main config override, merged over config.yaml. NOTE: the
# chart's own default for this key is REPLACED, not merged, so its
# auto_registration block is repeated verbatim below - drop it and the
# agent can no longer register itself.
config.yaml.local: |
api:
server:
auto_registration: # Activate if not using TLS for authentication
enabled: true
token: "${REGISTRATION_TOKEN}" # /!\ Do not modify this variable (auto-generated and handled by the chart)
allowed_ranges: # /!\ Make sure to adapt to the pod IP ranges used by your cluster
- "127.0.0.1/32"
- "192.168.0.0/16"
- "10.0.0.0/8"
- "172.16.0.0/12"
# This homelab has no egress to console.crowdsec.cloud: DNS does
# not resolve. The LAPI kept trying anyway ("Signal push: N
# signals to push", "capi metrics: sending" every 10s) and each
# attempt sat on a resolver timeout WHILE HOLDING A WRITE
# TRANSACTION, which is what kept stalling per-request decision
# lookups even with WAL enabled. Nothing to share and nothing to
# pull - turn the Central API off instead of letting it block the
# only database writer we have.
online_client:
sharing: false
pull:
community: false
blocklists: false
disable_usage_metrics_export: true
db_config:
# SQLite without WAL serialises every reader behind the writer's
# rollback journal, and the LAPI writes constantly: the agent pushes
# Traefik alerts read from Loki, the metrics collector counts
# decisions, the bouncer touches "last pull" on every request.
# Symptom: decision lookups taking 10-30s (and a second connection
# that could not even open the database) while the LAPI sat at 28m
# CPU - the process was blocked in fsync, not computing. Every
# bouncer-protected request then blew through the plugin timeout and
# fail-closed with 403, on every site at once.
# The PVC is local-path-retain (hostPath), not a network share, so
# WAL is safe here; the crowdsec docs recommend it for exactly this
# ("allowing more concurrency in SQLite that will improve
# performances in most scenarios").
use_wal: true
# Keeps the alert table bounded. At the 5000/7d default the file
# reached 54MB in 15 days off the Traefik access log alone, and the
# metrics collector counts decisions on a timer; a smaller working
# set means fewer full scans. Crowdsec only prunes - SQLite never
# shrinks the file, so the size stays until a manual VACUUM.
flush:
max_items: 1000
max_age: 24h
lapi: lapi:
env: env:
+8 -20
View File
@@ -30,10 +30,14 @@
# 3. ensure the static machine exists, recreating it with the # 3. ensure the static machine exists, recreating it with the
# Secret password if missing (agent retry loops reconnect # Secret password if missing (agent retry loops reconnect
# on their own - same name + same password); # on their own - same name + same password);
# 4. prune bouncer entries idle for 30d; # 4. prune bouncer entries idle for 30d.
# 5. delete decisions from LePresidente/http-generic-403-bf, a hub #
# scenario that bans an IP for 4h after 5 POST-403s in 10s and # It used to also delete LePresidente/http-generic-403-bf decisions hourly.
# therefore bans us for our own bouncer's fail-closed 403s. # That was a workaround for the bouncer failing closed on a slow LAPI and
# 403-ing the deploy runner into a 4h ban. The bouncer now polls decisions
# into a cache and never blocks on an unreachable LAPI, so it cannot
# manufacture those 403s any more, and the scenario only fires against real
# scanners - deleting their decisions hourly was undoing a working ban.
# #
# Manual apply (crowdsec/k8s is NOT managed by deploy.yaml): # Manual apply (crowdsec/k8s is NOT managed by deploy.yaml):
# kubectl apply -f crowdsec/k8s/janitor-cronjob.yaml # kubectl apply -f crowdsec/k8s/janitor-cronjob.yaml
@@ -196,19 +200,3 @@ spec:
fi fi
echo "== 4. prune stale bouncers (no pull for 30d) ==" echo "== 4. prune stale bouncers (no pull for 30d) =="
$LAPI_EXEC cscli bouncers prune -d 720h --force $LAPI_EXEC cscli bouncers prune -d 720h --force
echo "== 5. drop http-403-bf decisions (4h self-bans) =="
# `LePresidente/http-generic-403-bf` (hub item
# crowdsecurity/http-generic-bf v0.9) bans any source IP
# after 5 POSTs answered 403 within 10s, for 4h. That
# includes 403s this homelab generates ITSELF (any
# bouncer fail-closed, any app CSRF/rate-limit 403), and a
# 4h ban on the runner/home IP silently breaks deploys and
# browsing. The scenario cannot be removed per-scenario -
# it is baked into a hub item, and disabling the whole
# base-http-scenarios collection would drop ~40 useful
# detections. Instead we keep the detection and drop its
# decisions hourly; the LAN/home whitelists in
# crowdsec-values.yaml handle the legit sources, so this
# only ever hits real scanners (who are re-banned anyway).
$LAPI_EXEC cscli decisions delete \
--scenario LePresidente/http-generic-403-bf --all || true
+6 -9
View File
@@ -16,15 +16,12 @@ spec:
- name: gitea-service - name: gitea-service
port: 3000 port: 3000
# Registry route: NO crowdsec-bouncer. # Registry route: NO crowdsec-bouncer.
# The bouncer plugin does a blocking `GET /v1/decisions` to the LAPI on # A deploy burst (runner Action API polls, `docker manifest inspect` per
# *every* request. A deploy burst (runner Action API polls, `docker # own image, containerd pulls, smoke probes) fires hundreds of parallel
# manifest inspect` per own image, containerd pulls, smoke probes) fires # registry calls, and a ban on the runner breaks every later job. This
# hundreds of parallel registry calls; LAPI saturation pushed the lookup # route only serves authenticated OCI traffic - registry tokens and
# past the plugin timeout, and the bouncer fail-closed with 403 - which # basic-auth are already handled by gitea - and scanners get nothing
# containerd surfaces as ErrImagePull/ImagePullBackOff on the next pod. # useful from /v2, so there is no bruteforce surface to protect here.
# This route only serves authenticated OCI traffic (registry tokens,
# basic-auth already handled by gitea) and scanners get nothing useful
# from /v2, so there is no bruteforce surface to protect here.
- match: Host(`gcr.forust.xyz`) && PathPrefix(`/v2`) - match: Host(`gcr.forust.xyz`) && PathPrefix(`/v2`)
kind: Rule kind: Rule
services: services:
+9 -6
View File
@@ -7,12 +7,18 @@ spec:
entryPoints: entryPoints:
- websecure - websecure
routes: routes:
# NO crowdsec-bouncer on the API routes. These are the mesh client's own
# endpoints: gRPC-gateway management calls plus signal/relay long-polling,
# authenticated by NetBird's token rather than by a login form. A ban here
# is self-defeating - the client needs the mesh to reach anything else, so
# CrowdSec banning it locks the peer out of the network it needs to
# function. It also backfires: a banned peer keeps retrying, every retry
# is another 403, and LePresidente/http-generic-403-bf turns five 403s in
# ten seconds into a 4h ban, so one 403 loop kept re-arming the ban.
# netbird-local below has always been exempt; this makes prod match.
- match: Host(`nb.forust.xyz`) && (PathPrefix(`/signalexchange.SignalExchange/`) || PathPrefix(`/management.ManagementService/`) || PathPrefix(`/management.ProxyService/`)) - match: Host(`nb.forust.xyz`) && (PathPrefix(`/signalexchange.SignalExchange/`) || PathPrefix(`/management.ManagementService/`) || PathPrefix(`/management.ProxyService/`))
kind: Rule kind: Rule
priority: 100 priority: 100
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: netbird-server-service - name: netbird-server-service
port: 80 port: 80
@@ -20,9 +26,6 @@ spec:
- match: Host(`nb.forust.xyz`) && (PathPrefix(`/relay`) || PathPrefix(`/ws-proxy/`) || PathPrefix(`/api`) || PathPrefix(`/oauth2`)) - match: Host(`nb.forust.xyz`) && (PathPrefix(`/relay`) || PathPrefix(`/ws-proxy/`) || PathPrefix(`/api`) || PathPrefix(`/oauth2`))
kind: Rule kind: Rule
priority: 100 priority: 100
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: netbird-server-service - name: netbird-server-service
port: 80 port: 80