show-status.sh could never report failure. The ansible task that wraps it
('Execute show-status.sh and fail on failure') therefore always passed, on every
host, regardless of node state. Three bugs, each masking the next:
1. $? read too late. `code=0` sits between the sync-status.sh call and
`if [ $? -ne 0 ]`. A plain assignment succeeds and overwrites $? with 0, so
the condition was ALWAYS false and the else branch always taken.
2. The else branch was inverted. It is the sync-status-SUCCEEDED path, yet it set
`code=1; any_failure=true` — marking healthy nodes as failures.
3. any_failure could never propagate. check_sync_status runs backgrounded (`&`),
i.e. in a subshell, so `any_failure=true` inside it is discarded; and the
`wait "$pid"` loop threw away each job's exit status.
(3) hid (1) and (2): a script that believed every node had failed still exited 0,
so nobody saw it.
Fix: capture rc immediately; restore the intended logic (success => 0, syncing or
lagging => tolerated, anything else => failure); propagate failure in the PARENT
via `wait "$pid" || any_failure=true`, since the subshell cannot.
Verified with a stubbed sync-status.sh:
scenario before after
all online 0 0
one syncing 0 0 (tolerated)
one lagging 0 0 (tolerated)
one ERROR 0 1
ALL error 0 1
Behaviour for healthy fleets is unchanged; only genuine failures now surface.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fenced nodes are dropped from COMPOSE_FILE so show-status silently omitted
them - a fenced-but-running node looked 'gone'. rpc-update now writes
/root/rpc/.fenced (node | until | reason | owner per active window); print
it as a footer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A hung check-health.sh (aztec-testnet, looping on an unresponsive reference RPC)
blocked show-status.sh's parallel 'wait' for 3.5h, hanging the whole fleet
rpc-update and holding the deploy lock. Each curl was bounded (-m 3) and the
retry loop capped (3x), but the call itself wasn't time-bounded.
- sync-status.sh: wrap each check-health.sh call in 'timeout ${HC_TIMEOUT:-30}'
(-> exit 124 + 'timeout' status on overrun).
- show-status.sh: wrap the whole per-node sync-status.sh call in
'timeout ${SYNC_TIMEOUT:-60}' so the parallel wait can never block forever.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>