Commit Graph
311 Commits
Author SHA1 Message Date
admin f696efef11 Add: mount-guard.py - pre-deploy bind mount existence + Postgres empty-data heuristic check
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 23:03:11 -07:00
admin 2e889323ef Fix: correct immich.env-example paths to match real on-disk data (UPLOAD_LOCATION, BULK_UPLOAD_LOCATION, DB_DATA_LOCATION) after 2026-08-26 incident where generic template paths were wrong
ci/woodpecker/push/deploy Pipeline failed
2026-08-26 22:51:37 -07:00
admin b620682a41 Fix: envparse.py strip mode now also drops top-level 'name:' key that docker compose config emits but Swarm's stack deploy schema rejects
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 22:37:38 -07:00
admin 87a254c902 Fix: envparse.py strip mode now collapses long-form depends_on mapping to Swarm-compatible short-form list
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 22:35:53 -07:00
admin efd8caa218 Fix: invoke git-guard.sh via bash explicitly so tracked file mode bit doesn't matter
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 22:33:21 -07:00
admin 70814970c7 Harden .gitignore: broaden .env exclusion to *.env (with explicit exceptions for global.env and *.env-example templates), ignore stray .bak files
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 22:32:32 -07:00
admin 32e30f7e8b Add: call git-guard.sh at top of stack-deploy.sh to enforce sync before deploy
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 22:31:12 -07:00
admin 0bc117c868 Add: git-guard.sh - pre-deploy sync check to prevent local/Gitea divergence
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 22:30:47 -07:00
AVB 7b5273b4b4 Update ai/ai.yaml
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 21:36:48 -07:00
AVB 9d0e9d9b60 Update ai/ai.yaml
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 21:28:03 -07:00
AVB 08949ba3a4 Update ai/ai.yaml
ci/woodpecker/push/deploy Pipeline was successful
2026-08-26 21:24:17 -07:00
admin ef4bcfcb25 ai: trigger redeploy to apply $$ escaping fix (PR #8) to litellm secrets
ci/woodpecker/push/deploy Pipeline was successful
Comment-only change. Forces a real deploy of the ai stack now that
PR #8 (envparse.py $$ escaping) and PR #9 (deploy/ folder exclusion)
are both merged, so litellm picks up the correctly-escaped
LITELLM_MASTER_KEY/LITELLM_SALT_KEY instead of the truncated values
currently running (truncated at the first literal '$' due to the
envsubst+Compose double-interpolation bug fixed in #8).
2026-08-26 21:18:26 -07:00
AVB 808483181b Merge pull request 'fix: exclude deploy/ folder from changed-stack detection' (#9) from fix-exclude-deploy-folder into main
ci/woodpecker/push/deploy Pipeline was successful
Reviewed-on: #9
https://ai.bryanmail.net/c/17c61a7d-f6cb-41f7-bd68-c17880df3646
2026-08-26 21:13:40 -07:00
admin f6fc288aa6 fix: exclude deploy/ from changed-stack detection
grep -E '^[^/.][^/]*/' matches ANY non-dot top-level folder in the
changed-files list, including deploy/ -- the shared tooling folder
synced by every deploy, not a stack. A PR touching only
deploy/envparse.py caused stack-deploy.sh to be invoked with "deploy"
as a stack name, which correctly errored ("No main compose file... in
.../deploy") since deploy/ has no deploy.yaml.

No live service was affected (the error occurs before any redeploy
attempt), but it produced a confusing FAIL on an otherwise-correct
change (PR #8) and could mask a real failure in the noise.

Adds `| grep -v '^deploy$'` after the folder-name extraction in all 5
places this pattern appears (validate, provision-secrets x2, deploy,
verify). deploy/ is already unconditionally rsynced at the top of the
deploy step regardless of which stacks changed, so excluding it from
the stack list is safe -- it will still be synced, just never treated
as a deployable stack.
2026-08-26 21:10:59 -07:00
AVB 701d1289ea Merge pull request 'fix: prevent envsubst+Compose double-interpolation from truncating $ secrets' (#8) from fix-dollar-double-interpolation into main
ci/woodpecker/push/deploy Pipeline failed
Reviewed-on: #8
https://ai.bryanmail.net/c/17c61a7d-f6cb-41f7-bd68-c17880df3646
2026-08-26 21:05:37 -07:00
admin d209fc3222 fix: escape literal $ in .env values before envsubst (Pattern B stacks)
Root cause of the LITELLM_MASTER_KEY/LITELLM_SALT_KEY truncation
incident (2026-08-26): stack-deploy.sh's single-file deploy path is
`envsubst "$VARS" < stack.yaml | docker stack deploy -c - stack`.
envsubst embeds the raw secret value into the compose YAML text. If
that value contains a literal '$' followed by word chars, the
resulting YAML now contains what looks like a second variable
reference. `docker stack deploy -c -` runs Compose's own interpolation
pass on that text before creating the service, finds no such env var,
and silently substitutes empty string -- truncating the secret in the
running container with no error.

Confirmed: an 87-char LITELLM_MASTER_KEY arrived in the ai_litellm
container as 73 chars, silently, on a real deploy.

This is not specific to ai -- it affects every Pattern B stack (host
.env + envsubst, not native Docker secrets): maintenance, media,
unifi, guacamole, security, auth, traefik, meshcentral, ddm. Any of
them could have a '$'-containing value truncating right now without
detection, since the failure produces no warning.

Fix: escape every literal '$' as '$$' in export/export_merged (which
feed the `eval` that sets envsubst's actual source values), before
envsubst ever sees them. envsubst does not interpret '$' in replacement
text, so the doubled dollar survives envsubst intact; Compose's own
interpolation pass then consumes exactly one level of escaping,
landing on the correct single '$' with no leftover false variable
reference. vars/vars_merged (envsubst's allowlist string, unrelated to
values) are untouched.

NOT deployed/merged yet -- pending review. The currently-running
ai_litellm service still has the truncated keys and needs a fresh
`stack-deploy.sh ai` run after this merges to pick up the corrected
values.
2026-08-26 20:56:23 -07:00
admin d7fe56f6fe ai: add doc comment noting secrets are now provisioned via Woodpecker
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
No functional change -- this comment-only edit exists to trigger a real
deploy of the ai stack so provision-secrets' ai) case, and the AI_
var-name fix from PR #7, get exercised end-to-end for the first time.

Documents that MCPO_API_KEY is intentionally still manual/unmigrated.
2026-08-25 23:48:47 -07:00
AVB d1986678bc Merge pull request 'fix(ai): use AI_-prefixed AWS key names in ai) provisioning case' (#7) from fix-ai-aws-var-names into main
ci/woodpecker/push/deploy Pipeline was successful
Reviewed-on: #7
2026-08-25 23:45:27 -07:00
admin 6af2633936 fix(ai): write AI_-prefixed AWS key names to ai.env, not plain names
ai.yaml's litellm service (as of commit 37ed671a, "Change AWS keys to
use Woodpecker Secrets") references ${AI_AWS_ACCESS_KEY_ID} /
${AI_AWS_SECRET_ACCESS_KEY} and renders them into the container as
plain AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY. The ai) provisioning
case added in the earlier secrets-migration PR wrote the plain
(unprefixed) names into ai.env instead, which would leave
${AI_AWS_ACCESS_KEY_ID} unresolved at compose-render time (renders
empty) -- silently breaking Bedrock auth in litellm on the next ai
stack deploy.

Fixed both the grep -vE exclusion pattern and the two printf lines to
use the AI_-prefixed names. All other migrated vars in ai.yaml use
plain names and are unaffected.

No other changes in this file.
2026-08-25 23:42:39 -07:00
AVB 27df7cf58f Merge pull request 'HOTFIX: restore $${VAR} escaping dropped by AI secrets migration rewrite' (#6) from hotfix-dollar-escaping into main
ci/woodpecker/push/deploy Pipeline was successful
Reviewed-on: #6
2026-08-25 23:34:08 -07:00
admin 5188aef250 hotfix: restore $${VAR} double-dollar escaping for all secret-backed vars
The AI secrets migration PR (#4) was a full-file rewrite of deploy.yml.
That rewrite mechanically dropped one $ from EVERY $${VAR} occurrence in
the file, not just the new AI additions -- silently reverting all
pre-existing secret references (SWARM_MANAGER_IP, IMMICH_*, GIT_*,
POSTGRESQL_*, VAULTWARDEN_*, ENTERTAINMENT_*, etc.) to single-dollar
form. Per this file's own header comment, Woodpecker blanks single-dollar
braced refs at compile time since secrets aren't in that variable map --
this is the exact "SWARM_MANAGER_IP secret is empty" failure mode
documented above, and it fired immediately on the first push after #4
merged.

Impact: provision-secrets/deploy/verify all failed at their first
if-empty guard and exited before any ssh/scp/rsync ran. No live secret,
service, or deployed stack was touched -- this was a CI-only outage.

Fix: restored $${VAR} for every secret-backed reference throughout the
file. CI_PIPELINE_FILES / CI_COMMIT_MESSAGE stay single-dollar (correct
-- those are Woodpecker compile-time metadata, not secrets). The \$FILE
/ \$TMP backslash-escaping inside the ai) case's remote SSH command is
unrelated and was already correct (it protects those local-to-remote
vars from expanding before the SSH payload is sent).

This is a straight revert-of-the-regression -- no new secrets, no logic
changes beyond restoring the escaping.
2026-08-25 23:27:47 -07:00
admin d4797d4b4b docs: note ai stack secrets in SECRETS.md (trivial commit to force clean pipeline run)
ci/woodpecker/push/deploy Pipeline failed
No stack files changed -- this commit exists only to trigger a fresh
Woodpecker pipeline run against current main + current secrets, since
"Restart" on the prior failed run was replaying a stale snapshot from
before the ai_* secrets existed.
2026-08-25 23:19:14 -07:00
AVB 7d26f97a1e Merge pull request 'ai: migrate AWS/LiteLLM/OpenWebUI/OAuth secrets to Woodpecker' (#4) from ai-secrets-migration into main
Reviewed-on: #4
2026-08-25 23:01:09 -07:00
admin 2959721d3c ai: migrate AWS/LiteLLM/OpenWebUI/OAuth secrets from ai.env to Woodpecker secrets
Adds 8 new from_secret-backed env vars to provision-secrets and rewrites
the `ai)` case to do a targeted update of only those 8 keys in the
remote ai/ai.env via grep -v + printf (no sed, safe for values containing
/, $, &, etc). All other lines in ai.env (MCPO_API_KEY, OAUTH_CLIENT_ID,
WEBUI_URL, etc.) are left completely untouched -- MCPO_API_KEY migration
is deferred to a follow-up per plan, and this change never reads or
writes that value.

New secrets required in Woodpecker (Settings -> Secrets) before merge:
  ai_aws_access_key_id
  ai_aws_secret_access_key
  ai_litellm_master_key
  ai_litellm_salt_key
  ai_litellm_db_password
  ai_webui_secret_key
  ai_open_webui_database_url
  ai_oauth_client_secret
2026-08-25 22:40:05 -07:00
admin 52cecf2cf9 vaultwarden: rotate admin_token secret to v2 (hashed value)
ci/woodpecker/push/deploy Pipeline failed
2026-08-25 21:56:53 -07:00
AVB 37ed671aaa Change AWS keys to use Woodpecker Secrets
ci/woodpecker/push/deploy Pipeline was successful
2026-08-25 20:16:54 -07:00
AVB 1374a1dc1f Update open-webui to 0.11.1
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
2026-08-25 16:48:40 -07:00
AVB c9b53ec3f0 Update homeassistant/homeassistant.yaml
ci/woodpecker/push/deploy Pipeline was successful
2026-08-24 21:48:11 -07:00
admin 2b9209d9d1 Deprecate: hwaccel.transcoding.yml - functionality moved to main immich.yml
ci/woodpecker/push/deploy Pipeline failed
ci/woodpecker/cron/renovate Pipeline was successful
2026-08-19 23:03:53 -07:00
admin 4907039db9 Deprecate: hwaccel.ml.yml - functionality moved to main immich.yml
ci/woodpecker/push/deploy Pipeline was canceled
2026-08-19 23:03:52 -07:00
admin bef97dd908 Fix: Add /dev/dri to immich-server for quicksync transcoding support
ci/woodpecker/push/deploy Pipeline failed
2026-08-19 23:03:37 -07:00
admin 7af75c43ed Revert: Use volumes instead of devices for DDM compatibility with Docker Swarm
ci/woodpecker/push/deploy Pipeline failed
2026-08-19 23:02:06 -07:00
admin 472dd3a34a Revert: Use volumes instead of devices for DDM compatibility with Docker Swarm
ci/woodpecker/push/deploy Pipeline was canceled
2026-08-19 23:02:04 -07:00
admin d145aeeebe feat: add woodpecker-mcp.mjs server script for MCPO
ci/woodpecker/push/deploy Pipeline failed
ci/woodpecker/cron/renovate Pipeline was successful
2026-08-18 22:54:06 -07:00
admin 284bf98a6d Fix: Correct hwaccel.ml.yml to extend immich-machine-learning with devices, not create separate service
ci/woodpecker/push/deploy Pipeline failed
ci/woodpecker/cron/renovate Pipeline was successful
2026-08-18 07:51:58 -07:00
admin 623af9577c Fix: Correct hwaccel.transcoding.yml to extend immich-server with devices, not create separate service
ci/woodpecker/push/deploy Pipeline was canceled
2026-08-18 07:51:49 -07:00
AVB 268f81eb08 Update immich/immich.yml
ci/woodpecker/push/deploy Pipeline failed
2026-08-17 23:51:37 -07:00
AVB cef680c8df Update immich/immich.yml
ci/woodpecker/push/deploy Pipeline failed
2026-08-17 23:49:59 -07:00
AVB e641b9f14a Update immich/immich.yml
ci/woodpecker/push/deploy Pipeline failed
2026-08-17 23:47:35 -07:00
admin a0a1a87e04 Remove: Delete immich.env (should not be version controlled)
ci/woodpecker/push/deploy Pipeline failed
2026-08-17 23:38:07 -07:00
admin 0349a948e5 Add: Create immich.env-example with placeholder values
ci/woodpecker/push/deploy Pipeline failed
2026-08-17 23:37:41 -07:00
admin 38ace2a109 Add: Create immich.env with required environment variables
ci/woodpecker/push/deploy Pipeline failed
2026-08-17 23:35:18 -07:00
admin 6ccd4cc34a Fix: Add missing image to quicksync service in hwaccel.transcoding.yml
ci/woodpecker/push/deploy Pipeline was canceled
2026-08-17 23:35:18 -07:00
admin 47e9f95fcb Fix: Add missing image to openvino service in hwaccel.ml.yml
ci/woodpecker/push/deploy Pipeline was canceled
2026-08-17 23:35:17 -07:00
admin aa87f11588 cutover: fix missing PGPASSWORD on all remote psql -h calls (Phases 3/5/8)
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
Root cause of the "Canary row did not propagate" Phase 3 failure (run
#7): legacy's pg_hba.conf requires scram-sha-256 for any non-local
connection, and every remote `psql -h <patroni-node>` call in this
script had never supplied a password at all. This was masked until the
prior stderr-capture fix (run #7) surfaced the real error:
"fe_sendauth: no password supplied" on every single attempt.

patroni-0/patroni-1 use the SAME PGadmin superuser + SAME
postgresql_password secret as legacy itself (per
postgresql-ha-staging.yaml), so the fix reads that secret once via
`docker exec "$LEGACY_CID" cat /run/secrets/postgresql_password` early
in Phase 3, then passes it to every remote psql call via
`docker exec -e PGPASSWORD=...` (not spliced into the bash -c string,
to avoid quoting hazards).

Found and fixed the identical missing-password pattern in THREE
places, all with the same root cause:
- Phase 3: the canary-propagation SELECT (where it was first caught)
- Phase 5: the post-promotion pg_is_in_recovery() check
- Phase 8: the alias write + both leader/replica visibility checks

Added "2026 run #8" entry to the script's own header FIX LOG. Not yet
re-validated by a run reaching past Phase 3.

See ADR-0001 note, Session Update 11 (to be added).
2026-08-08 14:01:57 -07:00
admin dbabc69c2a cutover: surface stderr in Phase 3's canary-propagation check (diagnostic fix)
ci/woodpecker/push/deploy Pipeline was successful
Phase 3's "Canary row did not propagate to <replica> within 5s" failure
has now recurred twice (2026 run #4, root-caused as the Phase 2 lag-check
bug; and the run immediately after the Ceph IOPS fix, cause unconfirmed)
with genuinely healthy Phase 1/2 beforehand both times. The per-attempt
SELECT against the replica was discarding stderr entirely (2>/dev/null),
so a real connection/auth error and a genuine multi-second replication
delay were indistinguishable in the log — both just showed "row not
found".

Each of the 5 propagation-check attempts now captures stderr to
/tmp/cutover_phase3_attempt_<N>.stderr (mirrors the existing Phase 8
pattern) and echoes result+stderr into the log per-attempt. The final
trigger_rollback() message on failure includes the last non-empty
stderr seen directly in the FAIL line. The initial canary write's
stderr is also no longer discarded. No behavior change to timing/retry
counts — purely additive diagnostics.

See ADR-0001 note, Session Update 10 (to be added).
2026-08-08 13:07:53 -07:00
admin 2ed406db82 rollback: add full-session logging (writes to same dir as backup)
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
Persists this run's entire stdout/stderr transcript to BACKUP_DIR
(/volume1/SMB-docker/backup) — the same directory the pg_dumpall
backup lands in. Appends to an already-open CUTOVER_SESSION_LOG if
invoked as a child of cutover.sh (merging into that session's single
transcript); opens its own rollback-standalone-<TS>.log if run
standalone (including a manual run long after the fact, per this
script's own asymmetry warning). This is the single highest-value
place for a durable transcript in the whole suite, since rollback.sh
failing partway is the one scenario RUNBOOK.md flags as requiring
manual intervention. Also logs the transcript path in the grace-window
refusal message so it's not lost even in that failure mode.
Console/SSH output unchanged (tee mirrors to both).

See ADR-0001 note, Session Update 9.
2026-08-05 15:08:47 -07:00
admin 653f308623 cutover: add full-session logging (writes to same dir as backup)
ci/woodpecker/push/deploy Pipeline was successful
Persists the ENTIRE multi-phase transcript (Pre-Phase-0 through Phase
11, including Phase 0's preflight.sh output and any automatically
triggered rollback.sh output) to a single timestamped
cutover-session-<TS>.log in BACKUP_DIR (/volume1/SMB-docker/backup) —
the same directory the pg_dumpall backup lands in. Exports
CUTOVER_SESSION_LOG so child preflight.sh/rollback.sh invocations
append to the same file instead of opening their own. Adds an EXIT
trap that always announces final exit code + log path, specifically
so the one scenario RUNBOOK.md flags as needing manual intervention
(rollback.sh itself failing partway) is still fully investigable
after the fact even without a live terminal. Console/SSH output is
unchanged (tee mirrors to both).

See ADR-0001 note, Session Update 9.
2026-08-05 15:07:22 -07:00
admin fcbaa2cba3 cutover: add full-session logging to preflight.sh (writes to same dir as backup)
ci/woodpecker/push/deploy Pipeline was successful
Persists this run's entire stdout/stderr transcript to BACKUP_DIR
(/volume1/SMB-docker/backup), the SAME location as the pg_dumpall
backup file itself, so a failed run can be investigated later even
without a live terminal attached. Appends to an already-open
CUTOVER_SESSION_LOG if invoked as a child of cutover.sh (one merged
multi-phase transcript per session); opens its own
preflight-standalone-<TS>.log if run directly. Also logs the repo's
git HEAD at cutover/ for traceability, per the ADR-0001 note's
local-checkout-drift lesson (Session Update 8).

See ADR-0001 note, Session Update 9.
2026-08-05 15:03:14 -07:00
admin 82a8414ffa cutover.sh: interactive stale-state cleanup prompt before Phase 0
ci/woodpecker/push/deploy Pipeline was successful
Real incident: a run had both patroni-0 and patroni-1 stuck forever on
"waiting for standby_leader to bootstrap", never even attempting to race
for the role. Root cause was leftover etcd/patroni data on disk from a
prior interrupted run (operator stopped it short) — the postgresqlha
stack itself was gone, but etcd-1/2/3-data still had persisted raft state
including a real /service/postgres-ha/initialize key and old replication
slot records. Fresh Patroni nodes booting against that non-fresh etcd
correctly concluded the cluster already existed and deferred forever
waiting for a leader that could never appear, since nobody actually held
the lock. rollback.sh's own data-dir wipe only fires when it detects the
HA stack IS currently present (Case D/E) — if the stack was already gone
by the time cleanup ran, its Case A path never touches the data dirs,
leaving exactly this trap.

Added a "Pre-Phase-0" check that runs before preflight.sh's ~6+ minute
pg_dumpall: detects a still-present postgresqlha stack OR non-empty
etcd-*/patroni-*-data left over from a prior run, and — since this is
destructive and the operator explicitly wants this to be a deliberate
choice, not silent automatic cleanup — prompts interactively before
tearing down/wiping. Non-interactive sessions (no tty) hard-fail with a
clear message rather than guessing; AUTO_CLEANUP=yes in the environment
skips the prompt for deliberate unattended re-runs. Legacy production
data is never touched by any of this.
2026-08-05 13:37:15 -07:00