admin
e57f4b3592
cutover.sh: extend Phase 2 cascade-replica wait window to 35 min
...
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
Phase 2's cascade-replica bootstrap does its own full basebackup FROM the
new standby_leader (an extra hop beyond Phase 1's legacy-direct copy), so
per operator request its window is extended further than Phase 1's —
from 1200s (20 min, matched to observed 8-11 min legacy-direct timing) to
2100s (35 min), giving more margin for the additional hop. Phase 1's
window is intentionally left unchanged at 1200s since it already has
comfortable headroom against the timing we've actually observed for that
specific bootstrap path.
2026-08-04 00:48:32 -07:00
admin
29cd73eabc
cutover.sh: dynamic standby_leader detection + extended bootstrap wait windows
...
ci/woodpecker/push/deploy Pipeline was successful
Two real-run bugs found and fixed:
1. Phase 1 hardcoded patroni-1 as the expected standby_leader. The
bootstrap-race winner is actually nondeterministic (Patroni/etcd lock
race) — the original prod attempt had patroni-0 win it instead, which
the dry run never exercised. Fixed by polling BOTH patroni-0 and
patroni-1 each iteration and capturing whichever wins into
$LEADER_HOST, with the other becoming $REPLICA_HOST. Every later phase
(2,3,4,5,6,8,10) now references $LEADER_HOST/$REPLICA_HOST instead of
hardcoded hostnames.
2. Phase 1's wait window (240s) and Phase 2's (300s) were both far shorter
than the observed real basebackup duration for a ~42GB cluster
(8-11 minutes per prior dry-run/live polling). The role only flips to
standby_leader/replica AFTER the full copy completes, so both timeouts
could fire — and did — while a legitimate basebackup was still
in-progress, triggering a false-negative rollback. Both windows
extended to 1200s (20 min), with per-poll data-dir size logging
(timeout-guarded du -sh) so progress is observable instead of a silent
binary wait.
2026-08-04 00:40:20 -07:00
admin
9c1f31a70e
ADR-0001 rollback.sh: accept any 2xx in consumer health checks (204 from Woodpecker was a false-positive failure)
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
2026-08-02 22:15:15 -07:00
admin
325daca86b
ADR-0001 cutover.sh: fix Phase 1 node-label check (broken Go template on hyphenated label) and accept any 2xx in consumer health checks
ci/woodpecker/push/deploy Pipeline was successful
2026-08-02 22:14:17 -07:00
admin
066cbaf97f
ADR-0001 Phase 3: add rollback.sh (standalone, idempotent, grace-window guard against post-cutover data loss)
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
2026-08-02 17:38:50 -07:00
admin
1203630bb5
ADR-0001 Phase 3: add cutover.sh (11-phase scripted cutover with auto-rollback)
ci/woodpecker/push/deploy Pipeline was successful
2026-08-02 17:32:42 -07:00
admin
a4f5d54b82
ADR-0001 Phase 3: add postgresql-ha-final.yaml (Stage 3 - HAProxy alias handoff, no external port/Traefik yet)
ci/woodpecker/push/deploy Pipeline was successful
2026-08-02 17:28:26 -07:00
admin
01768d14f5
ADR-0001 Phase 3: add preflight.sh (disk/backup hard gates, catalog+consumer enumeration)
ci/woodpecker/push/deploy Pipeline was successful
2026-08-02 16:45:44 -07:00
admin
f49b576806
ADR-0001 Phase 3: add cutover RUNBOOK (dependency map, phased procedure, rollback)
ci/woodpecker/push/deploy Pipeline was successful
2026-08-02 16:35:52 -07:00
admin
8d30f19c8c
postgresql.yaml: update stale comment referencing old CLONE_* rationale to reference standby_cluster design (no functional change - this file never used CLONE_* itself)
ci/woodpecker/push/deploy Pipeline was successful
2026-07-30 22:03:44 -07:00
admin
4dcb3561ed
postgresql-ha-staging.yaml: switch from CLONE_WITH_BASEBACKUP to Patroni standby_cluster (continuous streaming, validated in pgha-test dry run) - closes pre-cutover write gap
ci/woodpecker/push/deploy Pipeline was successful
2026-07-30 21:59:44 -07:00
admin
dc6120ffd4
pgha-dryrun.yaml: trigger secret re-provisioning after regenerating postgresql_replication_password (excludes &<>\" per pystache HTML-escaping bug found this session)
ci/woodpecker/push/deploy Pipeline was successful
2026-07-30 21:44:49 -07:00
admin
8afe0049c4
pgha-dryrun.yaml: switch from CLONE_WITH_BASEBACKUP to Patroni standby_cluster (continuous streaming) to close the pre-cutover write gap
ci/woodpecker/push/deploy Pipeline was successful
2026-07-30 21:24:28 -07:00
admin
499402cd0d
postgresql-ha-staging.yaml: port dry-run fixes (bugs 1,2,4,5) - $$(...) escaping (incl. CLONE_PASSWORD), ETCD3_HOSTS, post_init_wrapper.sh SUPERUSER role fix
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
2026-07-30 15:56:43 -07:00
admin
7e8e3388f4
postgresql.yaml: port dry-run fixes (bugs 1,2,4,5) - $$(...) escaping, ETCD3_HOSTS, post_init_wrapper.sh SUPERUSER role fix
ci/woodpecker/push/deploy Pipeline was successful
2026-07-30 15:55:26 -07:00
admin
b7f0d9dbfb
pgha dry-run: fix bug 5 - grant SUPERUSER to placeholder postgres role
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
2026-07-29 09:32:09 -07:00
admin
e8e506dc4d
add: post_init_wrapper.sh to work around Spilo's hardcoded 'postgres' role name assumption in post_init.sh. Does not fork/modify Spilo's script — creates missing role then execs the original unchanged.
ci/woodpecker/push/deploy Pipeline was successful
2026-07-29 07:21:34 -07:00
admin
7529e6cb36
fix: override bootstrap.post_init via SPILO_CONFIGURATION to create missing 'postgres' role before Spilo's real post_init.sh runs. Spilo hardcodes ALTER VIEW...OWNER TO postgres with no way to parameterize, which fails since our superuser is PGadmin not postgres.
ci/woodpecker/push/deploy Pipeline was successful
2026-07-29 06:35:24 -07:00
admin
b86784fe3a
fix: legacy container needs pg_hba.conf replication rule for pg_basebackup — added initdb.d hook script. Disposable test only, uses 'trust' since this container is not auth-representative.
ci/woodpecker/push/deploy Pipeline was successful
2026-07-28 21:27:13 -07:00
admin
1dcb5ec467
fix: use ETCD3_HOSTS not ETCD_HOSTS — confirmed against spilo source that etcd/etcd3 are distinct DCS backends (v2 vs v3 API). Our etcd 3.5.9 containers have v2 API disabled, causing 404s with the old var name.
ci/woodpecker/push/deploy Pipeline was successful
2026-07-28 21:10:13 -07:00
admin
83480aa23b
fix: escape \$(...) as \$\$(...) in patroni command blocks — Compose interpolation was choking on \$( before the shell ever saw it (invalid interpolation format error)
ci/woodpecker/push/deploy Pipeline was successful
2026-07-28 21:03:56 -07:00
admin
20e1210441
postgresql: add disposable dry-run test stack for cutover validation (pgha-test, own network + data dirs, zero prod impact)
ci/woodpecker/push/deploy Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
2026-07-28 11:12:19 -07:00
admin
f4f0749969
postgresql: add cutover staging compose (ADR-0001 Phase 3). Lives in cutover/ subdir — deliberately excluded from stack-deploy.sh folder merge. Deployed only by cutover script as separate postgresqlha stack.
ci/woodpecker/push/deploy Pipeline was successful
2026-07-28 11:07:03 -07:00
admin
87ae54fb3f
postgresql: add HAProxy config for Patroni-aware TCP routing (ADR-0001 Phase 2)
ci/woodpecker/push/deploy Pipeline was successful
2026-07-28 08:53:55 -07:00
admin
c8d45e9e61
postgresql: add folder-based Patroni+etcd+HAProxy HA stack draft (ADR-0001 Phase 2). NOT deployed — bootstrap tier, manual deploy only. Coexists with flat postgresql.yaml until cutover.
ci/woodpecker/push/deploy Pipeline was canceled
2026-07-28 08:53:44 -07:00