Merge pull request 'traefik: add log rotation for access.log / traefik.log' (#23) from fix/traefik-log-rotation into main
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/push/deploy Pipeline was successful
Reviewed-on: #23 https://ai.bryanmail.net/c/5d375026-71a6-444a-a406-55fd55ad1f73
This commit was merged in pull request #23.
This commit is contained in:
@@ -0,0 +1,120 @@
|
|||||||
|
# Traefik Log Rotation
|
||||||
|
|
||||||
|
## Why this exists
|
||||||
|
|
||||||
|
`traefik/traefik.yaml` runs Traefik with:
|
||||||
|
- `--accesslog.filePath=/traefik/logs/access.log`
|
||||||
|
- `--log.filePath=/traefik/logs/traefik.log`
|
||||||
|
- `--log.level=DEBUG`
|
||||||
|
|
||||||
|
Neither file has any built-in rotation -- Traefik has no native rotate-on-size
|
||||||
|
nor a SIGUSR1/reopen handler. Docker's `json-file` log-driver rotation
|
||||||
|
(`max-size`/`max-file`) only applies to stdout, not to files Traefik writes
|
||||||
|
directly via `--accesslog.filePath`/`--log.filePath`. Result: `access.log` grew
|
||||||
|
to **~15.5GB unrotated** before this was caught, on a CephFS volume
|
||||||
|
(`/volume1/docker-root`) already at 82-84% used. Both the disk pressure and
|
||||||
|
the ongoing write latency of appending to a 15GB file on a network filesystem
|
||||||
|
on every request through Traefik were flagged as a real risk factor during
|
||||||
|
troubleshooting (Uptime Kuma WebSocket flapping investigation, Sep 2026).
|
||||||
|
|
||||||
|
**Note:** during that investigation, the actual root cause of the WebSocket
|
||||||
|
flapping turned out to be Uptime Kuma monitor misconfigurations (a Postgres
|
||||||
|
monitor throwing a null-reference error, and several monitors failing TLS
|
||||||
|
validation against self-signed/internal-IP certs) -- not this log file. This
|
||||||
|
rotation fix is still worth doing as general disk/IO hygiene, just not
|
||||||
|
causally tied to that incident.
|
||||||
|
|
||||||
|
## Architecture constraint this design accounts for
|
||||||
|
|
||||||
|
`traefik_reverse-proxy` runs Swarm *`mode: global`* -- one instance on **each**
|
||||||
|
of docker-1, docker-2, docker-3. All three write to the **same physical file**
|
||||||
|
via the shared CephFS bind mount `/volume1/docker/traefik -> /traefik`
|
||||||
|
(identical mount, visible identically from any node). That rules out:
|
||||||
|
|
||||||
|
- **Signal-based rotation** (classic `create` + `postrotate` sending SIGUSR1):
|
||||||
|
Traefik doesn't implement a reopen signal, and even if it did, you'd need to
|
||||||
|
signal 3 separate per-node containers in lockstep.
|
||||||
|
- **Running logrotate on just one node**: works until that node is down, then
|
||||||
|
rotation silently stops with no alert.
|
||||||
|
|
||||||
|
So this setup uses:
|
||||||
|
1. **`copytruncate`** (see `traefik-logs.conf`) -- all three Traefik processes
|
||||||
|
keep writing to the same inode, no signaling needed. Tradeoff: a few log
|
||||||
|
lines written in the exact copy/truncate instant can be lost -- fine for
|
||||||
|
diagnostic logs.
|
||||||
|
2. **Cron on all three nodes**, coordinated via a shared `flock` + shared
|
||||||
|
logrotate state file, both also on the CephFS mount (see
|
||||||
|
`traefik-logrotate.sh`). Whichever node's cron fires first grabs the lock,
|
||||||
|
rotates if due, and updates the shared state so the other two nodes' cron
|
||||||
|
runs see it's already done. No single node is a rotation SPOF.
|
||||||
|
3. **Size-triggered** (`size 250M`) rather than calendar-triggered (`daily`) --
|
||||||
|
this file can grow fast under bursts (see incident background above); a
|
||||||
|
purely daily interval would still let it balloon between runs. Cron checks
|
||||||
|
every 15 minutes, so it cannot grow much past the 250M threshold in practice.
|
||||||
|
|
||||||
|
## Install (one-time, per node)
|
||||||
|
|
||||||
|
Must run on **all three** docker LXCs -- this is host-level cron/logrotate
|
||||||
|
config, not something `stack-deploy.sh` can reach (it only touches Swarm
|
||||||
|
services, not host cron jobs).
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# On each of docker-1 (192.168.4.31), docker-2 (192.168.4.32), docker-3 (192.168.4.33):
|
||||||
|
ssh root@<192.168.4.31|.32|.33>
|
||||||
|
cd /volume1/docker/compose-files
|
||||||
|
bash deploy/git-guard.sh # confirm sync first, as always
|
||||||
|
sudo bash traefik/logrotate/install.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
Verify on each node:
|
||||||
|
```bash
|
||||||
|
cat /etc/cron.d/traefik-logrotate
|
||||||
|
logrotate -d /etc/logrotate.d/traefik-logs # dry-run, confirms syntax
|
||||||
|
```
|
||||||
|
|
||||||
|
## Bootstrapping -- shrinking the *existing* oversized log file
|
||||||
|
|
||||||
|
**Not done automatically by `install.sh`.** The new size-triggered config only
|
||||||
|
prevents *future* unbounded growth -- it won't touch the current 15.5GB file
|
||||||
|
until the next time it crosses 250M (i.e. never, since it's already well past
|
||||||
|
that and logrotate only acts on crossing the threshold going forward from its
|
||||||
|
recorded size at last check).
|
||||||
|
|
||||||
|
`copytruncate` always copies the full current file before truncating it --
|
||||||
|
that's inherent to how it works, not a bug. Forcing a rotation of the current
|
||||||
|
15.5GB file would momentarily need roughly another 15-GB-sized chunk of free
|
||||||
|
space -- and `/volume1/docker-root` only had **~14G free** at last check.
|
||||||
|
Doing this blindly could tip the volume to 100% mid-operation.
|
||||||
|
|
||||||
|
**Recommended manual step (operator-run, on any one node -- it's the same
|
||||||
|
CephFS file from all three)**: these are diagnostic access/debug logs, not
|
||||||
|
something worth preserving in full, so just truncate directly rather than
|
||||||
|
compress-then-truncate:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Optional: keep a small tail sample for reference before truncating
|
||||||
|
tail -c 50000000 /volume1/docker/traefik/logs/access.log > \
|
||||||
|
/volume1/docker/traefik-access-log-archive-$(date +%Y%m%d).log
|
||||||
|
|
||||||
|
# Then truncate in place (safe -- Traefik's existing file handles on all 3
|
||||||
|
# nodes stay valid, same as copytruncate's own mechanism):
|
||||||
|
: > /volume1/docker/traefik/logs/access.log
|
||||||
|
: > /volume1/docker/traefik/logs/traefik.log
|
||||||
|
|
||||||
|
# Confirm:
|
||||||
|
df -h /volume1/docker-root
|
||||||
|
ls -la /volume1/docker/traefik/logs/
|
||||||
|
```
|
||||||
|
|
||||||
|
After this one-time bootstrap, the cron+logrotate setup keeps it bounded
|
||||||
|
(rotates at 250M, keeps 48 compressed generations, prunes anything over 14
|
||||||
|
days old) going forward without needing any further manual intervention.
|
||||||
|
|
||||||
|
## Related, NOT included in this change (follow-up to consider separately)
|
||||||
|
|
||||||
|
`traefik.yaml` currently runs `--log.level=DEBUG` -- verbose debug logging in
|
||||||
|
production, which is a meaningful contributor to how fast these files grow.
|
||||||
|
Lowering to `INFO` would reduce volume significantly but requires a real
|
||||||
|
modification + redeploy of the high-blast-radius `traefik` stack (all
|
||||||
|
HTTP/HTTPS routing depends on it), so it's intentionally left out of this PR
|
||||||
|
and should be its own reviewed change if wanted.
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# One-time installer for Traefik log rotation.
|
||||||
|
#
|
||||||
|
# MUST be run manually, once, on EACH of docker-1, docker-2, docker-3.
|
||||||
|
# This is intentionally NOT part of stack-deploy.sh / the Woodpecker
|
||||||
|
# pipeline: it installs a host-level cron.d entry and /etc/logrotate.d
|
||||||
|
# config, and `docker stack deploy` has no mechanism to reach outside the
|
||||||
|
# Swarm/container boundary onto host cron. See README.md for why this needs
|
||||||
|
# to run on all three nodes.
|
||||||
|
#
|
||||||
|
# Usage (from a checkout of this repo, on each node):
|
||||||
|
# sudo bash traefik/logrotate/install.sh
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
if [[ $EUID -ne 0 ]]; then
|
||||||
|
echo "Run as root (sudo)." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
REPO_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
|
||||||
|
echo "Installing logrotate config..."
|
||||||
|
install -m 0644 "$REPO_DIR/traefik-logs.conf" /etc/logrotate.d/traefik-logs
|
||||||
|
|
||||||
|
echo "Installing rotation wrapper..."
|
||||||
|
install -m 0755 "$REPO_DIR/traefik-logrotate.sh" /usr/local/sbin/traefik-logrotate.sh
|
||||||
|
|
||||||
|
echo "Installing cron.d schedule (every 15 minutes)..."
|
||||||
|
cat > /etc/cron.d/traefik-logrotate << 'EOF'
|
||||||
|
# Managed by homelab/compose-files traefik/logrotate/install.sh -- do not
|
||||||
|
# hand-edit; update traefik/logrotate/*.{conf,sh} in Gitea and re-run
|
||||||
|
# install.sh instead.
|
||||||
|
*/15 * * * * root /usr/local/sbin/traefik-logrotate.sh
|
||||||
|
EOF
|
||||||
|
chmod 0644 /etc/cron.d/traefik-logrotate
|
||||||
|
|
||||||
|
echo "Verifying logrotate config syntax..."
|
||||||
|
logrotate -d /etc/logrotate.d/traefik-logs
|
||||||
|
|
||||||
|
echo "Done. First rotation check runs on the next cron tick (up to 15 min)."
|
||||||
|
echo "This only rotates going forward once the file crosses the size threshold."
|
||||||
|
echo "To shrink the EXISTING already-large log file, see README.md -- that is"
|
||||||
|
echo "a separate, deliberate manual step (not done by this script) because it"
|
||||||
|
echo "needs disk headroom awareness first."
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Wrapper invoked by cron on docker-1/docker-2/docker-3 to rotate the shared
|
||||||
|
# Traefik access/error logs (see traefik-logs.conf for the "full why").
|
||||||
|
#
|
||||||
|
# Because /volume1/docker/traefik/logs is the SAME physical CephFS path on
|
||||||
|
# all three nodes, and cron on all three nodes runs this independently, we
|
||||||
|
# use a shared flock (also on the CephFS mount, so it's visible cluster-wide)
|
||||||
|
# to guarantee only one node actually executes logrotate at a time, and a
|
||||||
|
# SHARED state file so whichever node runs it knows the true last-rotated
|
||||||
|
# time regardless of which node rotated it last. If a node is down, the
|
||||||
|
# other two still cover the schedule -- none of this relies on a specific
|
||||||
|
# node being up.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
LOCK_DIR="/volume1/docker/traefik/logrotate-state"
|
||||||
|
LOCK_FILE="$LOCK_DIR/rotate.lock"
|
||||||
|
STATE_FILE="$LOCK_DIR/status"
|
||||||
|
CONF_FILE="/etc/logrotate.d/traefik-logs"
|
||||||
|
|
||||||
|
mkdir -p "$LOCK_DIR"
|
||||||
|
touch "$STATE_FILE"
|
||||||
|
|
||||||
|
exec 200>"$LOCK_FILE"
|
||||||
|
if ! flock -n 200; then
|
||||||
|
# Another node already holds the lock this cycle -- normal, not an error.
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
/usr/sbin/logrotate -s "$STATE_FILE" "$CONF_FILE"
|
||||||
@@ -0,0 +1,43 @@
|
|||||||
|
# Traefik access/error log rotation
|
||||||
|
#
|
||||||
|
# CONTEXT: traefik_reverse-proxy runs in Swarm `mode: global` (traefik.yaml),
|
||||||
|
# meaning one instance runs on EACH of docker-1/docker-2/docker-3. All three
|
||||||
|
# write to the SAME physical file via the shared CephFS bind mount
|
||||||
|
# /volume1/docker/traefik/logs -> /traefik (identical path from any node --
|
||||||
|
# see infra context: /volume1/docker is a shared CephFS mount).
|
||||||
|
#
|
||||||
|
# `copytruncate` is REQUIRED here (not the default create+signal approach):
|
||||||
|
# Traefik has no SIGUSR1/SIGHUP "reopen log file" handling, and even if it
|
||||||
|
# did, coordinating a reopen signal across 3 independent per-node containers
|
||||||
|
# writing to one shared inode is unnecessary complexity. copytruncate keeps
|
||||||
|
# every writer's existing file descriptor valid (truncates in place) so all
|
||||||
|
# three Traefik processes keep appending to the same inode with zero
|
||||||
|
# signaling. Tradeoff: a handful of log lines written in the exact
|
||||||
|
# copy/truncate instant can be lost -- acceptable for diagnostic access/error
|
||||||
|
# logs, not used for anything transactional.
|
||||||
|
#
|
||||||
|
# Size-triggered (not calendar-triggered) on purpose: this file can grow fast
|
||||||
|
# under bursts (see incident that prompted this -- 15.5GB accumulated with
|
||||||
|
# --log.level=DEBUG set). `size` is checked every time the wrapper script runs
|
||||||
|
# (cron, every 15 minutes -- see install.sh), so it cannot balloon unbounded
|
||||||
|
# between checks the way a plain `daily` interval would.
|
||||||
|
#
|
||||||
|
# Installed via install.sh on ALL THREE docker LXCs (docker-1, docker-2,
|
||||||
|
# docker-3) -- see README.md. This is a HOST-level cron/logrotate config,
|
||||||
|
# outside the Woodpecker/stack-deploy.sh pipeline (docker stack deploy has no
|
||||||
|
# mechanism to touch host cron), so it must be applied manually once per node,
|
||||||
|
# not via a stack redeploy.
|
||||||
|
|
||||||
|
/volume1/docker/traefik/logs/*.log {
|
||||||
|
size 250M
|
||||||
|
rotate 48
|
||||||
|
maxage 14
|
||||||
|
compress
|
||||||
|
delaycompress
|
||||||
|
missingok
|
||||||
|
notifempty
|
||||||
|
copytruncate
|
||||||
|
dateext
|
||||||
|
dateformat -%Y%m%d-%H%M%S
|
||||||
|
su root root
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user