Altinity / Altinity/altinity-sql-browser
clickhouse-containers.mjs: attachDockerNetworkWithRollback races container cold-start instead of confirmed readiness
Nobody has claimed this yet.
- Dominant language
- TypeScript
- Stars
- 8
- Forks
- 2
- Avg merge
- 1h 34m
- Merged PRs (30d)
- 6
Description
attachDockerNetworkWithRollback's docstring in tests/spike/clickhouse-client/clickhouse-containers.mjs:292-297 claims the port was "confirmed reachable at container-boot time by the caller before this runs" — but startRow calls it at line 419 before waitForReady at line 422. Nothing confirms reachability before the attach runs.
The function's 4×1s probePing (lines 277-289, 306) therefore races ClickHouse's own cold start rather than checking an already-confirmed-live port. Found live and reproducibly (3/3 runs) during #585 Phase 0 WebKit-browser-matrix flake research: current-altinity-stable is consistently the row that loses this race and gets silently rolled back to default-bridge-only networking, because it's booted last (sequential boot order in spike-server.mjs:167-173) under maximum accumulated Docker load — so its cold start is slowest and most likely to still be starting when the probe fires.
This is comment/invariant drift (the docstring asserts a precondition the call site doesn't actually provide) that makes container network topology depend on relative boot speed rather than a real readiness check. Low urgency — this harness is dev/spike-only, not production — but worth fixing before the harness is relied on again for a rerun of the #585 browser matrix, since it's a plausible contributor to that matrix's one flaky cell (see #585 ship-log / ADR-0005 evidence discussion).
Suggested fix: either call attachDockerNetworkWithRollback only after waitForReady resolves (matching the docstring's own claimed precondition), or have the docstring/precondition match reality (loosen probePing's retry budget, or make it wait for the same readiness signal waitForReady uses).
Found by: automated root-cause research launched from a /ship-adjacent session, 2026-08-06.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start in tests/spike/clickhouse-client/clickhouse-containers.mjs at attachDockerNetworkWithRollback (lines 277-306) and startRow (lines 419-422), then compare its probePing behavior with waitForReady. Check spike-server.mjs:167-173 for the sequential boot order and reproduce the slow-start case. Done means the documented readiness invariant is true at the call site and the container network is not silently rolled back during cold start.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- docker, javascript
- Domain
- infrastructure, testing
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 68/100