traefik: add log rotation for access.log / traefik.log #23

Merged
AVB merged 9 commits from fix/traefik-log-rotation into main 2026-09-12 00:29:28 -07:00
4 changed files with 236 additions and 0 deletions
+120
View File
@@ -0,0 +1,120 @@
# Traefik Log Rotation
## Why this exists
`traefik/traefik.yaml` runs Traefik with:
- `--accesslog.filePath=/traefik/logs/access.log`
- `--log.filePath=/traefik/logs/traefik.log`
- `--log.level=DEBUG`
Neither file has any built-in rotation -- Traefik has no native rotate-on-size
nor a SIGUSR1/reopen handler. Docker's `json-file` log-driver rotation
(`max-size`/`max-file`) only applies to stdout, not to files Traefik writes
directly via `--accesslog.filePath`/`--log.filePath`. Result: `access.log` grew
to **~15.5GB unrotated** before this was caught, on a CephFS volume
(`/volume1/docker-root`) already at 82-84% used. Both the disk pressure and
the ongoing write latency of appending to a 15GB file on a network filesystem
on every request through Traefik were flagged as a real risk factor during
troubleshooting (Uptime Kuma WebSocket flapping investigation, Sep 2026).
**Note:** during that investigation, the actual root cause of the WebSocket
flapping turned out to be Uptime Kuma monitor misconfigurations (a Postgres
monitor throwing a null-reference error, and several monitors failing TLS
validation against self-signed/internal-IP certs) -- not this log file. This
rotation fix is still worth doing as general disk/IO hygiene, just not
causally tied to that incident.
## Architecture constraint this design accounts for
`traefik_reverse-proxy` runs Swarm *`mode: global`* -- one instance on **each**
of docker-1, docker-2, docker-3. All three write to the **same physical file**
via the shared CephFS bind mount `/volume1/docker/traefik -> /traefik`
(identical mount, visible identically from any node). That rules out:
- **Signal-based rotation** (classic `create` + `postrotate` sending SIGUSR1):
Traefik doesn't implement a reopen signal, and even if it did, you'd need to
signal 3 separate per-node containers in lockstep.
- **Running logrotate on just one node**: works until that node is down, then
rotation silently stops with no alert.
So this setup uses:
1. **`copytruncate`** (see `traefik-logs.conf`) -- all three Traefik processes
keep writing to the same inode, no signaling needed. Tradeoff: a few log
lines written in the exact copy/truncate instant can be lost -- fine for
diagnostic logs.
2. **Cron on all three nodes**, coordinated via a shared `flock` + shared
logrotate state file, both also on the CephFS mount (see
`traefik-logrotate.sh`). Whichever node's cron fires first grabs the lock,
rotates if due, and updates the shared state so the other two nodes' cron
runs see it's already done. No single node is a rotation SPOF.
3. **Size-triggered** (`size 250M`) rather than calendar-triggered (`daily`) --
this file can grow fast under bursts (see incident background above); a
purely daily interval would still let it balloon between runs. Cron checks
every 15 minutes, so it cannot grow much past the 250M threshold in practice.
## Install (one-time, per node)
Must run on **all three** docker LXCs -- this is host-level cron/logrotate
config, not something `stack-deploy.sh` can reach (it only touches Swarm
services, not host cron jobs).
```bash
# On each of docker-1 (192.168.4.31), docker-2 (192.168.4.32), docker-3 (192.168.4.33):
ssh root@<192.168.4.31|.32|.33>
cd /volume1/docker/compose-files
bash deploy/git-guard.sh # confirm sync first, as always
sudo bash traefik/logrotate/install.sh
```
Verify on each node:
```bash
cat /etc/cron.d/traefik-logrotate
logrotate -d /etc/logrotate.d/traefik-logs # dry-run, confirms syntax
```
## Bootstrapping -- shrinking the *existing* oversized log file
**Not done automatically by `install.sh`.** The new size-triggered config only
prevents *future* unbounded growth -- it won't touch the current 15.5GB file
until the next time it crosses 250M (i.e. never, since it's already well past
that and logrotate only acts on crossing the threshold going forward from its
recorded size at last check).
`copytruncate` always copies the full current file before truncating it --
that's inherent to how it works, not a bug. Forcing a rotation of the current
15.5GB file would momentarily need roughly another 15-GB-sized chunk of free
space -- and `/volume1/docker-root` only had **~14G free** at last check.
Doing this blindly could tip the volume to 100% mid-operation.
**Recommended manual step (operator-run, on any one node -- it's the same
CephFS file from all three)**: these are diagnostic access/debug logs, not
something worth preserving in full, so just truncate directly rather than
compress-then-truncate:
```bash
# Optional: keep a small tail sample for reference before truncating
tail -c 50000000 /volume1/docker/traefik/logs/access.log > \
/volume1/docker/traefik-access-log-archive-$(date +%Y%m%d).log
# Then truncate in place (safe -- Traefik's existing file handles on all 3
# nodes stay valid, same as copytruncate's own mechanism):
: > /volume1/docker/traefik/logs/access.log
: > /volume1/docker/traefik/logs/traefik.log
# Confirm:
df -h /volume1/docker-root
ls -la /volume1/docker/traefik/logs/
```
After this one-time bootstrap, the cron+logrotate setup keeps it bounded
(rotates at 250M, keeps 48 compressed generations, prunes anything over 14
days old) going forward without needing any further manual intervention.
## Related, NOT included in this change (follow-up to consider separately)
`traefik.yaml` currently runs `--log.level=DEBUG` -- verbose debug logging in
production, which is a meaningful contributor to how fast these files grow.
Lowering to `INFO` would reduce volume significantly but requires a real
modification + redeploy of the high-blast-radius `traefik` stack (all
HTTP/HTTPS routing depends on it), so it's intentionally left out of this PR
and should be its own reviewed change if wanted.
+44
View File
@@ -0,0 +1,44 @@
#!/usr/bin/env bash
# One-time installer for Traefik log rotation.
#
# MUST be run manually, once, on EACH of docker-1, docker-2, docker-3.
# This is intentionally NOT part of stack-deploy.sh / the Woodpecker
# pipeline: it installs a host-level cron.d entry and /etc/logrotate.d
# config, and `docker stack deploy` has no mechanism to reach outside the
# Swarm/container boundary onto host cron. See README.md for why this needs
# to run on all three nodes.
#
# Usage (from a checkout of this repo, on each node):
# sudo bash traefik/logrotate/install.sh
set -euo pipefail
if [[ $EUID -ne 0 ]]; then
echo "Run as root (sudo)." >&2
exit 1
fi
REPO_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
echo "Installing logrotate config..."
install -m 0644 "$REPO_DIR/traefik-logs.conf" /etc/logrotate.d/traefik-logs
echo "Installing rotation wrapper..."
install -m 0755 "$REPO_DIR/traefik-logrotate.sh" /usr/local/sbin/traefik-logrotate.sh
echo "Installing cron.d schedule (every 15 minutes)..."
cat > /etc/cron.d/traefik-logrotate << 'EOF'
# Managed by homelab/compose-files traefik/logrotate/install.sh -- do not
# hand-edit; update traefik/logrotate/*.{conf,sh} in Gitea and re-run
# install.sh instead.
*/15 * * * * root /usr/local/sbin/traefik-logrotate.sh
EOF
chmod 0644 /etc/cron.d/traefik-logrotate
echo "Verifying logrotate config syntax..."
logrotate -d /etc/logrotate.d/traefik-logs
echo "Done. First rotation check runs on the next cron tick (up to 15 min)."
echo "This only rotates going forward once the file crosses the size threshold."
echo "To shrink the EXISTING already-large log file, see README.md -- that is"
echo "a separate, deliberate manual step (not done by this script) because it"
echo "needs disk headroom awareness first."
+29
View File
@@ -0,0 +1,29 @@
#!/usr/bin/env bash
# Wrapper invoked by cron on docker-1/docker-2/docker-3 to rotate the shared
# Traefik access/error logs (see traefik-logs.conf for the "full why").
#
# Because /volume1/docker/traefik/logs is the SAME physical CephFS path on
# all three nodes, and cron on all three nodes runs this independently, we
# use a shared flock (also on the CephFS mount, so it's visible cluster-wide)
# to guarantee only one node actually executes logrotate at a time, and a
# SHARED state file so whichever node runs it knows the true last-rotated
# time regardless of which node rotated it last. If a node is down, the
# other two still cover the schedule -- none of this relies on a specific
# node being up.
set -euo pipefail
LOCK_DIR="/volume1/docker/traefik/logrotate-state"
LOCK_FILE="$LOCK_DIR/rotate.lock"
STATE_FILE="$LOCK_DIR/status"
CONF_FILE="/etc/logrotate.d/traefik-logs"
mkdir -p "$LOCK_DIR"
touch "$STATE_FILE"
exec 200>"$LOCK_FILE"
if ! flock -n 200; then
# Another node already holds the lock this cycle -- normal, not an error.
exit 0
fi
/usr/sbin/logrotate -s "$STATE_FILE" "$CONF_FILE"
+43
View File
@@ -0,0 +1,43 @@
# Traefik access/error log rotation
#
# CONTEXT: traefik_reverse-proxy runs in Swarm `mode: global` (traefik.yaml),
# meaning one instance runs on EACH of docker-1/docker-2/docker-3. All three
# write to the SAME physical file via the shared CephFS bind mount
# /volume1/docker/traefik/logs -> /traefik (identical path from any node --
# see infra context: /volume1/docker is a shared CephFS mount).
#
# `copytruncate` is REQUIRED here (not the default create+signal approach):
# Traefik has no SIGUSR1/SIGHUP "reopen log file" handling, and even if it
# did, coordinating a reopen signal across 3 independent per-node containers
# writing to one shared inode is unnecessary complexity. copytruncate keeps
# every writer's existing file descriptor valid (truncates in place) so all
# three Traefik processes keep appending to the same inode with zero
# signaling. Tradeoff: a handful of log lines written in the exact
# copy/truncate instant can be lost -- acceptable for diagnostic access/error
# logs, not used for anything transactional.
#
# Size-triggered (not calendar-triggered) on purpose: this file can grow fast
# under bursts (see incident that prompted this -- 15.5GB accumulated with
# --log.level=DEBUG set). `size` is checked every time the wrapper script runs
# (cron, every 15 minutes -- see install.sh), so it cannot balloon unbounded
# between checks the way a plain `daily` interval would.
#
# Installed via install.sh on ALL THREE docker LXCs (docker-1, docker-2,
# docker-3) -- see README.md. This is a HOST-level cron/logrotate config,
# outside the Woodpecker/stack-deploy.sh pipeline (docker stack deploy has no
# mechanism to touch host cron), so it must be applied manually once per node,
# not via a stack redeploy.
/volume1/docker/traefik/logs/*.log {
size 250M
rotate 48
maxage 14
compress
delaycompress
missingok
notifempty
copytruncate
dateext
dateformat -%Y%m%d-%H%M%S
su root root
}