Files
homelab/netbird
forust 892790822d fix(crowdsec): stop the 403 loop that banned our own VPS and runner
Chasing why apply-k8s kept dying mid-run turned up a self-inflicted
ban loop. 585 of 586 LePresidente/http-generic-403-bf alerts in the LAPI
came from 193.181.211.79 - our own VPS - POSTing
/management.ManagementService/GetServerKey, i.e. the NetBird client's
own management call. netbird-server had never seen a single one of them,
so the 403 was not NetBird's: the bouncer was rejecting the request
before it got there. A banned peer keeps retrying, each retry is another
403, and the scenario turns five 403s in ten seconds into a 4h ban, so
the loop kept re-arming the ban it was serving. The hourly janitor step
that deleted those decisions hourly was masking all of it.

* netbird/k8s/ingress.yaml - drop the bouncer from the mesh API routes
  (gRPC-gateway management, signal, relay, /api, /oauth2). A ban there
  locks a peer out of the network it needs to reach anything else, and
  those endpoints authenticate by NetBird token, not by a login form.
  The dashboard keeps the bouncer; it is a real login surface.
  netbird-local was already exempt, so this makes prod match.

* crowdsec-middleware.yaml - CrowdsecMode stream instead of live. live
  blocked on GET /v1/decisions per request, so a burst saturated the
  LAPI and the plugin 403'd IPs that were never banned. v1.3.3 ignores
  UpdateMaxFailure in live, so fail-open is only reachable in stream;
  -1 now means an unreachable LAPI degrades to "no protection" rather
  than "every site 403". 15s poll instead of the 60s default, because
  the runner shares one public IP with the house.

  Also corrects the key name: HTTPTimeoutSeconds, not
  CrowdsecLapiTimeout, which never existed and was being silently
  dropped, leaving the 10s default. Back at 10, not the 2s f7cd75d
  guessed - a pull that times out leaves the ban cache frozen at its
  startup contents, so new bans would silently never apply.

* crowdsec-values.yaml - CIDR allowlisting moves to parsers/s02-enrich,
  where CrowdSec's docs put it: a parser whitelist drops the event before
  it reaches a bucket, so those addresses never become a decision at
  all. The old postoverflow LAN list was checked only after the ban
  existed, which is the window the deploy kept landing in. Added
  100.64.0.0/10, which the RFC 1918 blocks miss and where the
  workstation, the k0s node and the VPS actually live. The DDNS
  home-IP whitelist stays in postoverflows, because resolving a hostname
  is the expensive check the docs reserve that stage for.

  ClientTrustedIPs mirrors that list so the bouncer skips the LAPI
  round-trip entirely for those addresses.

* janitor-cronjob.yaml - drop step 5. The bouncer can no longer
  manufacture 403s, so the only remaining firings of that scenario are
  real scanners, and deleting their decisions hourly was undoing a
  working ban.

* Also lands the LAPI config.yaml.local (SQLite WAL, Central API off,
  bounded flush) that f7cd75d's comments referenced but never included:
  the LAPI was blocked in fsync on its rollback journal, and the CAPI
  resolver held a write transaction while timing out against a host
  this network cannot reach.

Verified in-cluster: no new 403-bf alerts in the 3.5min after applying,
the management endpoint answers 404 from the backend instead of 403 from
the bouncer in 0.17s, and the VPS client reports Management and Signal
connected with 2/2 relays.
2026-09-27 18:09:07 +02:00
..

NetBird

Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind the repository's existing external Traefik instance on the Docker proxy network. Only STUN UDP 3478 is published directly.

The deployment uses SQLite for a single-instance homelab server. The persistent netbird_data volume and the datastore encryption key are both required to recover the installation.

Files

  • compose.yaml: dashboard and combined server; selected by the marker-driven deploy workflow through active.
  • config.template.yaml: non-secret server configuration rendered at startup.
  • entrypoint.sh: injects Docker secrets into an in-memory runtime configuration.
  • client.compose.yaml: optional host-network peer using a dashboard-generated setup key.
  • .env: ignored local hostnames, the detected Traefik Docker-network subnet, and optional client setup key.
  • secrets/: ignored relay secret and datastore encryption key.

First deployment

Run these commands on the Docker host before merging the activating branch. The deploy preflight resets tracked files but preserves ignored local state.

cd /srv/homelab/netbird
./setup.sh
$EDITOR .env
docker compose config --quiet
docker compose up -d

Review the values in .env before starting. The example public hostname is netbird.forust.xyz; change it if a different public domain was selected. setup.sh replaces NETBIRD_PROXY_SUBNET=auto with the first IPv4 subnet of the external Docker proxy network. Keep that value synchronized with the network; set an explicit CIDR instead if the network is managed elsewhere.

setup.sh is idempotent and never replaces existing secrets. Do not delete or regenerate secrets/datastore-encryption-key after the first successful start unless all encrypted setup keys and API tokens are intentionally being invalidated.

Network prerequisites

  • Point the public hostname directly to the Docker host. Do not proxy UDP 3478 through Cloudflare or another CDN.
  • Allow inbound TCP 80, TCP 443, and UDP 3478 through the host firewall and upstream router.
  • Ensure the external proxy Docker network exists and Traefik uses its websecure entrypoint and letsencrypt resolver. NETBIRD_PROXY_SUBNET must describe that network; it is used to trust only forwarded client addresses from Traefik.
  • Ensure the internal names in .env resolve where the local and development aliases are needed.
  • Keep Traefik's websecure read timeout disabled for long-lived gRPC and WebSocket sessions. This repository configures --entrypoints.websecure.transport.respondingTimeouts.readTimeout=0 in traefik/compose.yaml.

After startup, verify OIDC discovery through the public TLS endpoint:

curl -fsS "https://${NETBIRD_DOMAIN}/oauth2/.well-known/openid-configuration"

Open https://${NETBIRD_DOMAIN} immediately and complete the initial owner setup. Treat the initial setup flow as public until the owner exists.

Optional host client

The client intentionally lives in a separate Compose project. Normal server deploys use --remove-orphans, so keeping the client in the server project would cause it to be removed.

  1. Create a reusable or ephemeral setup key in the NetBird dashboard.
  2. Put NB_SETUP_KEY=<key> in the ignored netbird/.env file.
  3. Set NETBIRD_CLIENT_HOSTNAME to this machine's desired peer name.
  4. Start and inspect the client:
cd /srv/homelab/netbird
docker compose -f client.compose.yaml config --quiet
docker compose -f client.compose.yaml up -d
docker compose -f client.compose.yaml exec netbird-client netbird status

The client uses host networking and requires NET_ADMIN, SYS_ADMIN, SYS_RESOURCE, and /dev/net/tun. Remove it without affecting the server stack:

docker compose -f client.compose.yaml down

Operations

Inspect status and logs:

docker compose ps
docker compose logs --tail=200 netbird-server dashboard

Stop or remove containers without deleting data:

docker compose down

Do not add -v to docker compose down; it would delete the NetBird datastore.

Backup and restore

Back up both the persistent volume and the ignored secret files. For a consistent SQLite backup, briefly stop the server first and store the resulting archive and datastore-encryption-key in an encrypted backup:

cd /srv/homelab/netbird
mkdir -p backups
docker compose stop netbird-server
docker run --rm \
  -v netbird_data:/data:ro \
  -v "$PWD/backups:/backup" \
  busybox:1.37.0 \
  tar -C /data -czf "/backup/netbird-data-$(date -u +%Y%m%dT%H%M%SZ).tar.gz" .
docker compose start netbird-server

Also securely back up:

  • secrets/datastore-encryption-key — required to decrypt stored secrets.
  • secrets/relay-auth-secret — keeps issued relay credentials valid across restoration.
  • netbird/.env — optional, but it records the public and internal hostnames.

Test a restore in an isolated Docker host before relying on a backup.

Upgrade

  1. Take and verify a backup.
  2. Review NetBird release notes for server, client, and dashboard compatibility.
  3. Update the pinned tags in compose.yaml; update client.compose.yaml separately when deploying the client.
  4. Pull and recreate the selected services:
docker compose pull
docker compose up -d

The image tags are intentionally pinned instead of using latest, matching this repository's pull-on-deploy policy.