From 2d9c2be9f30365491daf38a9e82ea4aa8c5d84a6 Mon Sep 17 00:00:00 2001 From: pingqiu Date: Sun, 26 Apr 2026 10:32:10 -0700 Subject: [PATCH] =?UTF-8?q?G5-4=20m01+M02=20cluster=20bring-up=20=E2=80=94?= =?UTF-8?q?=20hand-off=20to=20sw?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Records QA's cross-node smoke attempt 2026-04-26: infrastructure fully verified READY (m01+M02 reachability, SMB share for binary distribution, master cross-node listen, network OK), but cluster bring-up blocked at V3-internal gate. Symptom: blockvolume on both nodes connects to master but logs "durable open: frontend: volume not ready" — never reaches steady state, status endpoint never binds, master log shows no heartbeat or assignment-mint events. Hand-off contents: - §1 specific questions for sw (5 gaps to fill) - §2 infrastructure verified READY (no action needed) - §3 copy-pasteable commands sw can run/debug (build → topology → master → primary → replica → cleanup) - §4 QA's hypothesis on the gap (assignment-from-master flow) - §5 debug suggestions for sw (log levels, integration test references) - §6 G5-4 script skeleton current state - §7 QA's next steps once sw answers Working dirs reproducible: - Binaries: /mnt/smb/work/share/g5-binaries/{blockmaster,blockvolume} - Run state: /tmp/g5sm/ on both nodes - Logs: /tmp/g5sm/logs/{master,primary,replica}.log Blocks: G5-4 implementation work (script scenario bodies, hardware first-light scenarios). Does NOT block QA scenario authoring at component scope (Cluster framework already covers that). Co-Authored-By: Claude Opus 4.7 (1M context) --- .../v3-phase-15-g5-m0102-bringup-handoff.md | 248 ++++++++++++++++++ 1 file changed, 248 insertions(+) create mode 100644 sw-block/design/v3-phase-15-g5-m0102-bringup-handoff.md diff --git a/sw-block/design/v3-phase-15-g5-m0102-bringup-handoff.md b/sw-block/design/v3-phase-15-g5-m0102-bringup-handoff.md new file mode 100644 index 000000000..d64781a3c --- /dev/null +++ b/sw-block/design/v3-phase-15-g5-m0102-bringup-handoff.md @@ -0,0 +1,248 @@ +# G5-4 m01+M02 Cluster Bring-Up — Hand-off to sw + +**Date**: 2026-04-26 +**Status**: ⏸ blocked on V3-internal bring-up sequence question +**From**: QA (round 2026-04-26 cross-node smoke attempt) +**To**: sw (G5-4 framework owner) +**Context**: G5-4 m01 hardware first-light per [g5-kickoff §3 batch G5-4](v3-phase-15-g5-kickoff.md). Skeleton script committed at `seaweed_block@eabafe8` (`scripts/iterate-m01-replicated-write.sh`). + +--- + +## §1 What I'm asking sw to answer + +**Question**: what's the canonical V3 flow to bring a 2-node cluster from cold-start to "primary + replica both healthy"? + +My attempt below got both volumes stuck at `blockvolume: durable open: frontend: volume not ready`. The volumes connect to master successfully but never reach "ready" state. + +**Specific gaps I need filled:** +1. Is there a missing CLI flag or config beyond what's listed in `--help`? +2. Does `topology.yaml` need fields beyond `volumes/slots/{replica_id,server_id}`? +3. Does master need an explicit "mint assignment" trigger, or does it fire automatically from topology + observed heartbeats? +4. Is there a settling period > 4 seconds expected before "ready"? +5. Is there example bring-up test code I can reference (e.g., sparrow integration test, or `cmd/blockmaster/*_test.go`)? + +--- + +## §2 Infrastructure verified READY (no action needed) + +| Layer | Status | How verified | +|---|---|---| +| m01 + M02 reachability | ✅ | `ping 192.168.1.184` from m01 = 0.92ms | +| SMB share cross-node binary distribution | ✅ | `v:/share` on Windows = `/mnt/smb/work/share/` on both Linux nodes | +| Binary execution on M02 | ✅ | M02 ran `blockvolume --help` from SMB share without rebuild | +| Master cross-node listen | ✅ | `blockmaster --listen 0.0.0.0:9180` bound; `ss -tlnp` confirms | +| Network reachability primary↔master, replica↔master | ✅ | Both `blockvolume` processes connected to master without error | + +--- + +## §3 What I ran (copy-paste reproducible) + +### 3.1 Build binaries on m01 + drop to SMB share + +```bash +ssh -i /c/work/dev_server/testdev_key testdev@192.168.1.181 " +cd /opt/work/seaweed_block_t4d4 && \ +go build -o /mnt/smb/work/share/g5-binaries/blockvolume ./cmd/blockvolume/ && \ +go build -o /mnt/smb/work/share/g5-binaries/blockmaster ./cmd/blockmaster/ +" +``` + +**Result**: ✅ both binaries built successfully (16 MiB blockmaster, 18 MiB blockvolume) + +### 3.2 Verify M02 can execute the binary + +```bash +ssh -i /c/work/dev_server/testdev_key testdev@192.168.1.184 \ + "/mnt/smb/work/share/g5-binaries/blockvolume --help 2>&1 | head -3" +``` + +**Result**: ✅ `Usage of blockvolume: -ctrl-addr string ...` (executes from SMB share without scp) + +### 3.3 Setup directories + topology YAML on m01 + +```bash +ssh -i /c/work/dev_server/testdev_key testdev@192.168.1.181 \ + 'mkdir -p /tmp/g5sm/{master-store,primary-durable,logs}' + +ssh -i /c/work/dev_server/testdev_key testdev@192.168.1.181 \ + "printf 'volumes:\n - volume_id: v1\n slots:\n - replica_id: r1\n server_id: m01-primary\n - replica_id: r2\n server_id: m02-replica\n' > /tmp/g5sm/topology.yaml && cat /tmp/g5sm/topology.yaml" +``` + +**Result**: ✅ topology.yaml created; schema deduced from `cmd/blockmaster/topology.go:30-40` struct tags + +```yaml +volumes: + - volume_id: v1 + slots: + - replica_id: r1 + server_id: m01-primary + - replica_id: r2 + server_id: m02-replica +``` + +### 3.4 Start blockmaster on m01 + +```bash +ssh -i /c/work/dev_server/testdev_key testdev@192.168.1.181 \ + "nohup /mnt/smb/work/share/g5-binaries/blockmaster \ + --authority-store /tmp/g5sm/master-store \ + --listen 0.0.0.0:9180 \ + --topology /tmp/g5sm/topology.yaml \ + --t0-print-ready \ + > /tmp/g5sm/logs/master.log 2>&1 /tmp/g5sm/logs/primary.log 2>&1 /tmp/g5sm/logs/replica.log 2>&1 /dev/null; sudo pkill -9 -f blockvolume 2>/dev/null" +ssh -i /c/work/dev_server/testdev_key testdev@192.168.1.184 \ + "sudo pkill -9 -f blockvolume 2>/dev/null" +``` + +**Result**: ✅ both nodes clean + +--- + +## §4 What I think the gap is (sw to confirm or correct) + +The error `blockvolume: durable open: frontend: volume not ready` happens BEFORE the status endpoint binds, BEFORE durable storage opens. The blockvolume seems to be waiting for an assignment-from-master before completing initialization. + +If that's correct, then either: +1. **Master needs to actively mint + push** assignment to volumes (not just have topology loaded passively), and there's a trigger I'm missing +2. **Volume needs to wait long enough** for master heartbeat → topology resolution → assignment dispatch (4s wasn't enough, but how long is right?) +3. **Topology YAML needs more** (e.g., `expected_servers` section, `epoch`, `endpoint_version`, or other authority fields) +4. **There's a bootstrap admin command** to trigger initial assignment dispatch + +A quick way to find out: sw can point me at any working bring-up integration test (probably in `core/replication/integration_*_test.go` or `cmd/blockmaster/*_test.go`) that brings a multi-node cluster up. I'll mirror its pattern in the script. + +--- + +## §5 What sw can do to debug + +### 5.1 Run the same sequence on m01 with fresher eyes + +All commands in §3 are copy-pasteable. The setup is reproducible: +- Binaries at `/mnt/smb/work/share/g5-binaries/{blockmaster,blockvolume}` (built 2026-04-26 from `seaweed_block@e642ae8`+ working tree at part C land time; rebuild if needed for `a0be6d5` test fixture race fix) +- Working dirs at `/tmp/g5sm/` on both m01 and M02 + +### 5.2 Possible things to try + +- Add `--log-level=debug` or similar verbose flag if blockvolume / blockmaster supports it +- Read `/tmp/g5sm/logs/{master,primary,replica}.log` after bring-up attempt — the silent failure suggests a log channel that's not flushing or not at default level +- Check if blockmaster needs explicit `slots[].expected` or `slots[].epoch` fields +- Check if there's a `blockadmin` CLI tool for triggering assignments + +### 5.3 If a working bring-up sequence is available somewhere + +Pointer to: +- An L3 integration test that spins up multi-node cluster +- An existing m01 script that brings up cluster (none in `seaweed_block/scripts/` other than my skeleton) +- `cmd/sparrow` test code (sparrow has integration tests; might include cluster bootstrap) +- Any documentation of the expected bring-up sequence + +Will let me update the G5-4 script skeleton with the right pattern. + +--- + +## §6 Where the G5-4 script skeleton lives + +`seaweed_block@eabafe8` — `scripts/iterate-m01-replicated-write.sh` (272 LOC). Marked DRAFT v0.1. + +Sections that work today (per §2 infra-verified): +- Config block (env-overridable) +- `sync_and_build` (build on m01, scp/SMB-share binary to M02) +- Helpers (log/die/collect_diagnostics) + +Section blocked on this hand-off: +- `start_cluster` — currently has the same flags I used in §3.4–3.6 above; will fail same way until bring-up sequence is correct +- All scenario bodies (TODO-marked, depend on `start_cluster` working) + +--- + +## §7 Once sw answers — QA next steps + +1. Update `iterate-m01-replicated-write.sh` `start_cluster` with correct bring-up sequence +2. Verify the corrected sequence brings cluster to "primary + replica healthy" state +3. Author scenario bodies per G5-4 architect ratification (currently awaiting in `g5-kickoff.md` §8) +4. Run full matrix: walstore + smartwal × 4 scenarios = 8 runs + +ETA after sw answers: ~half day to update + verify; scenario bodies depend on G5 mini-plan.