A1 Engine kind-routing fix: SessionProgressObserved/Completed/Failed now respect active session Kind. Rebuild progress no longer leaks into catch-up aggregate. sessionKindMismatch guard + observeRebuildProgress helper. 2 regression tests lock kind isolation. A2 Retention pin: Rebuild session ack drives progress-based WAL retention floor. Pin installed at base_lsn on accepted, advances with wal_applied_lsn, released on completed/failed/cancelled. rebuildProgressPinFloor returns min across all active replicas. Retention pin test: 100 blocks fill WAL, 5 flusher cycles with 20 pinned rebuild entries — all verified correct. A3 Progress ack emission: Automatic sessionAck(running/base_complete/completed/failed) emitted from rebuild session lifecycle transitions. sessionAckLocked builds ack under session lock. emitRebuildSessionAck callback wired through SetOnRebuildSessionAck on BlockVol. ObserveReplicaRebuildSessionAck maps acks to core engine events. WireLocalReplicaRebuildSessionAcks bridges local callback to server. 5 server tests proving ack→core, pin advance, pin cleanup. A4 Deadline/timeout: rebuildAckWatch watchdog: armed on accepted/running/base_complete, refreshed on each ack, cleared on completed/failed. Timeout cancels local session + clears pin + fail-closes. 2 tests: timeout→fail-close, progress→refresh. A5 Session-controlled execution path: v2bridge.Executor.TransferFullBase now uses session-controlled loop: beginControlledFullBase → real sessionControl over TCP → transferExtentToSession via RebuildTransportClient → PrepareFullBaseRebuild → TryCompleteRebuildSession. ReplicaReceiver control channel handles MsgSessionControl alongside MsgBarrierReq. Session acks written back on same TCP connection. RebuildSessionBase request type separates new per-block stream from legacy raw extent stream. Full-base cleanup deferred until success. Deadlock fix: ApplyBaseBlock releases session lock before ioMu. Hydration skip for full-base sessions. 23 rebuild component tests (all pass): 11 kernel correctness, 8 transport/runtime, 3 scenario-scale, including 1GB primary-initiated with CRC validation. 29 files changed, ~2500 insertions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
6.8 KiB
V2 Kernel Closure Review
Date: 2026-04-05 Status: active
Question
The goal is not to prove whether iSCSI itself can be implemented. The reusable
blockvol + frontend code already shows that.
The real question is whether the current kernel split can grow into a product:
- Is the brain owned by V2 semantics?
- Is the control plane owned by V2 messages and convergence?
- Is the data plane attached as an execution/backend service instead of a truth owner?
Current Answer
The current masterv2 + volumev2 + purev2 shape is viable as a product kernel
because ownership is split in the right direction.
Brain
Owner:
sw-block/engine/replication/
What it owns:
- semantic state
- event ingestion
- command intent
- outward projection
What it must not own:
- backend I/O
- transport lifecycle
- frontend serving details
Control Plane
Owner:
sw-block/runtime/masterv2/sw-block/runtime/volumev2/control_session.gosw-block/runtime/volumev2/orchestrator.go
What it owns:
- desired declaration
- heartbeat observation
- assignment emission
- assignment apply loop
- convergence/idempotence
What it may normalize but must not own semantically:
- raw
syncAck/ timeout / callback observations may be normalized into primary-side sync fact kinds such as:sync_quorum_ackedsync_quorum_timed_outsync_replay_requiredsync_rebuild_requiredsync_replay_failed
- this normalization is only a control-plane envelope for "what fact arrived"
- it must not become a second decision authority or a second session planner
What it must not own:
- WAL/extent execution
- frontend protocol implementation
Data Plane
Owner:
sw-block/runtime/purev2/sw-block/runtime/volumev2/frontend.go- reused
weed/storage/blockvol/*
What it owns:
- create/open
- read/write/flush
- restart durability
- frontend export such as iSCSI
What it must not own:
- role truth
- publication truth
- assignment policy
Closure Proofs
Two small closure proofs are enough for the current stage.
1. Control-plane closure
Scenario:
masterv2declares one RF1 primaryvolumev2heartbeats with no local role yetmasterv2emits an assignmentvolumev2applies it through the V2 path- a later heartbeat converges to quiet state
- if desired state changes, assignment is reissued once and converges again
Why it matters:
- this proves the new head is not piggybacking on
weed/serverloops
2. Data-plane closure
Scenario:
volumev2exports a named volume through iSCSI- a client logs in and issues SCSI write/read
- data is verified through the frontend and local backend view
Why it matters:
- this proves the kernel can host a real frontend while keeping truth ownership outside the frontend/backend code
Product Meaning
If these two closures stay true while features expand, then the architecture can scale toward:
- RF1 productized single-node block service
- RF2/RF3 replication as additional control/data workflows
- failover and rebuild without moving semantic truth back into backend code
- CSI on top of a clearer runtime contract
Another way to state the same result:
masterv2behaves like an external identity authority- each
volumev2instance behaves like a per-volume micro-cluster shell - the selected primary inside that shell owns data-control truth and recovery choreography
Current Milestone
The current milestone is:
- live transport-backed failover-time evidence now crosses one real loopback HTTP path
- one continuous Loop 2 service and one bounded auto-failover service now exist
- one runtime-managed frontend path and one bounded repair/catch-up wrapper now exist
- one end-to-end RF2 handoff proof now exists with continued I/O on the new primary
- one bounded operator surface and one bounded CSI runtime backend adapter now sit on top of runtime-owned truth
What this milestone proves:
- the new kernel can now carry RF2 failover, active replication observation, bounded continuity, real serving, and one outward/operator/CSI surface stack without collapsing authority ownership
- runtime/product-facing state can now be attached as compressed projection rather than by inventing a new truth owner
- the
masterv2identity boundary and primary-led data-control boundary still remain intact - one bounded working RF2 block path now exists
What it does not yet prove:
- broad RF2 product approval across deployment and frontend/operator matrices
- full rebuild lifecycle choreography beyond the bounded repair wrapper
- multi-process / multi-host proof for the current working path
- pilot-ready or launch-ready working block behavior
Next Major Milestone
The next major milestone should be:
multi-process and pilot-ready RF2 validation
This means one level above the current runtime/product surface slice:
- prove the current working path outside the current bounded runtime harness
- widen the working path into multi-process or multi-host validation
- harden rebuild lifecycle and operator/CSI behavior on that wider path
- attach pilot/preflight/containment evidence on top of the widened path
Target Shape
Code should look roughly like:
masterv2: identity authority onlyvolumev2: runtime-owned failover, active service Loop 2, repair, continuity, and projected RF2 surfacespurev2: execution adapter and local boundary observation- transport/session adapters: live participant communication beyond the current bounded harness
- product-facing layer: bounded frontend/operator/CSI attachment with pilot artifacts
Exit Criteria
The milestone should be considered complete when all are true:
- the current bounded working path is proven outside the current harness
- pilot/preflight/containment evidence exists for that wider path
- no new product/runtime surface silently widens into broad launch approval
Non-Goals
This milestone should not try to prove:
- broad product launch approval
- broad matrix approval across all transport/frontend combinations
- broad
RF>2product closure - reopening the kernel authority split
It should prove only that the current working RF2 block path can widen toward pilot-ready validation without collapsing the ownership model.
Main Risk
The main risk is not iSCSI or local I/O. The main risk is semantic leakage:
- adding more backend-state shortcuts into control decisions
- letting frontend/backend code redefine publication truth
- rebuilding
weed/serverownership insidevolumev2 - letting
masterv2grow from identity authority into a centralized recovery planner - letting normalized sync facts drift into a hidden second protocol vocabulary with no explicit review surface
As long as those risks are resisted, the kernel can keep expanding cleanly.