mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-09-20 13:30:46 +02:00
Add host-side protocol state seam that derives per-replica execution state from V2 sender/session snapshots and blocks live-tail WAL shipping while an active recovery session is in progress. New file: weed/server/block_protocol_state.go - replicaProtocolExecutionState derived from engine snapshots - LiveEligible=false during active catch-up/rebuild sessions - bindProtocolExecutionPolicy wires policy into BlockVol - syncProtocolExecutionState called after assignments + core events Data plane changes: - WALShipper.Ship() checks liveShippingPolicy before dial/send - BlockVol.SetLiveShippingPolicy persists across shipper group rebuilds - ShipperGroup propagates policy to all shippers Design contract: sw-block/design/v2-protocol-aware-execution.md Scope: WAL-first rollout only. Prevents illegal live-tail delivery during active recovery. Does not change snapshot/build behavior or move backlog. Next wave: bounded WAL catch-up under same contract. Tests: 4 unit/component tests for phase gate behavior, plus bootstrap seam tests that confirmed the two pre-existing bugs locally. 13 files changed, 900 insertions, 69 deletions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
3.2 KiB
3.2 KiB
V2 Protocol-Aware Execution
Purpose
Make host-side execution in weed/server and weed/storage/blockvol obey the
existing V2 session contract explicitly. The engine remains the semantic source
of truth. Host code owns only:
- execution-state caching derived from sender/session snapshots
- phase gating before data-plane I/O
- observation routing back into core events
Host-Side Execution State
For each primary volume and replica, the host caches a replica protocol execution state with these fields:
ReplicaIDSenderStateSessionIDSessionKindSessionPhaseStartLSNTargetLSNFrozenTargetLSNRecoveredToSessionActiveLiveEligibleReason
Rules:
- State is derived from
v2Orchestrator.Registrysnapshots only. LiveEligible=falsewhenever there is an active recovery session.- Data-plane code must consult this cached state before shipping current live WAL entries.
- Heartbeat and publication remain projection-driven; they do not invent local session semantics.
WAL-First Rollout
The first rollout is intentionally narrow:
- cover
keepupand WAL-based catch-up only - do not change snapshot/build policy
- do not let fresh late-attached replicas consume current live-tail WAL while a bounded catch-up session is active
Current implementation seam:
weed/server/block_protocol_state.go- derives host execution state from sender/session snapshots
- binds a per-volume live-shipping policy back into
BlockVol
weed/storage/blockvol/blockvol.go- carries the host-provided live-shipping policy across shipper-group rebuilds
weed/storage/blockvol/wal_shipper.go- checks the policy before any live-tail dial or send
This is intentionally a phase gate, not a second source of truth.
Observation Seam
Runtime observations should feed back through one server-side seam:
- sender/session snapshots ->
syncProtocolExecutionState() - host event application ->
applyCoreEvent() - assignment processing ->
ApplyAssignments()
The rule is:
- engine chooses the protocol phase
- host derives execution state from engine snapshots
- data path obeys that state
- host emits observed facts back through
applyCoreEvent()
Fast Test Roster
The first fast-test roster for protocol-aware execution is:
unit:TestWALShipper_LiveShippingPolicyBlocksBeforeDial- proves phase gate happens before any transport dial
unit:TestWALShipper_LiveShippingPolicyAllowsShip- proves the gate does not block normal live shipping after eligibility
component:TestBlockService_ProtocolExecutionState_ActiveCatchUpBlocksLiveShipping- proves sender/session snapshots become host execution state and block live shipping during active catch-up
component:TestBlockService_ProtocolExecutionState_InSyncSenderAllowsLiveShipping- proves the host reopens live shipping after the recovery session is gone
Next fast tests to add in later waves:
- late attach with backlog must stay bounded until target reached
- transport contact before barrier durability must not imply publish healthy
- timeout with valid retention pin may replan WAL catch-up
- timeout after retention loss must escalate to build