diff --git a/sw-block/.private/phase/phase-20-acceptance.md b/sw-block/.private/phase/phase-20-acceptance.md index 96f1c37d3..e62bcb08f 100644 --- a/sw-block/.private/phase/phase-20-acceptance.md +++ b/sw-block/.private/phase/phase-20-acceptance.md @@ -56,7 +56,9 @@ Current `Phase 20` reading: 1. the hard-blocker closure set is `Implemented` 2. the hard-blocker closure set is `Developer-validated` on the bounded acceptance subset 3. the hard-blocker closure set is not yet globally `Tester-validated` just because the developer proof passed -4. tester automation now has a metadata-driven suite entry for `Stage 0`, but the current hardware run still fails at the known first-write `dd_write` / `sync_all` barrier issue, so acceptance remains pending +4. tester automation now has metadata-driven suite entries for both `Stage 0` and `Stage 1` +5. `Stage 0` bootstrap closure is now proven on real hosts: `create -> 10s wait -> 4k fsync -> publish_healthy` +6. the remaining hardware failure has been isolated to `Stage 1` sustained workload under the default `64MB` WAL budget, so overall acceptance still remains pending ## Tester Validation Still Required @@ -79,8 +81,9 @@ only "currently believed" but regression-frozen. Current tester status: 1. the metadata-driven suite pipeline now runs end-to-end: build, deploy, remote scenario execution, and evidence collection -2. `P20-H0` currently starts and runs remotely, but the first hardware run still fails in `record-before` on the known first-write `dd_write` path -3. this means tester infrastructure is now real and reusable, but `Stage 0` is not yet a passing acceptance artifact +2. `P20-H0` is now a passing hardware artifact for the bounded bootstrap claim and should be treated as the `Stage 0` closure case +3. the failing `record-before` workload has been moved conceptually into `Stage 1`, where it now reads as a WAL-budget / sustained-I/O issue rather than a bootstrap protocol gap +4. this means tester infrastructure is real and reusable, `Stage 0` is closed, and the next hardware blocker is the master-managed WAL-size gap for `Stage 1` This checklist is intentionally concrete. Each row should answer: diff --git a/sw-block/.private/phase/phase-20-t6-runbook.md b/sw-block/.private/phase/phase-20-t6-runbook.md index 8615e5d58..197619824 100644 --- a/sw-block/.private/phase/phase-20-t6-runbook.md +++ b/sw-block/.private/phase/phase-20-t6-runbook.md @@ -85,7 +85,8 @@ No promotion-mode toggle required. Goal: -1. close the bootstrap membership gap on real hosts +1. prove the bootstrap membership gap is closed on real hosts +2. freeze `create -> first fsync fence -> publish_healthy` as the standalone `P20-H0` artifact #### Stage 1 @@ -127,10 +128,10 @@ Keep the toggle visible in scenario source or in a wrapper-generated temp copy. ## Stage 0 Bootstrap Closure Checklist -`Stage 0` is the current hard gate before meaningful `V2` failover hardware -validation. +`Stage 0` is now the bounded bootstrap artifact that must stay green before +interpreting broader failover runs. -The blocker observed on hardware is: +The blocker that originally motivated this checklist was: 1. promoted primary still shows `ReplicaIDs=[]` 2. `RoleApplied=true` @@ -210,6 +211,12 @@ Healthy RF2 path must show all of the following: 6. master `cluster_replication_mode` returns to a healthy cluster judgment 7. no persistent `projection_mismatches` remain for the healthy path +Current reading: + +1. the dedicated bootstrap-only scenario now passes on hardware +2. `P20-H0` should therefore be treated as the closed `Stage 0` case +3. sustained `fio + dd_write` failure after bootstrap belongs to `Stage 1`, not to this checklist + ### Stage 0 Fail Criteria Any one of the following keeps `Stage 0` open: @@ -225,7 +232,7 @@ Any one of the following keeps `Stage 0` open: ### Minimum Operator Loop -When iterating on the fix, record this sequence each run: +When validating or rechecking the bootstrap closure, record this sequence each run: 1. before failure: `block/volume/` and `/debug/block/shipper` 2. immediately after failover: same two surfaces @@ -256,6 +263,7 @@ Stage 1 means: 2. do not enable `--block.v2Promotion=true` 3. capture `block/volume/` before and after failover 4. capture `/debug/block/shipper` on both candidate servers during the run +5. read sustained post-bootstrap write failures as `Stage 1` workload issues unless bootstrap itself regresses ### Stage 1 Must Prove @@ -301,6 +309,12 @@ One scenario at a time: sw-test-runner run weed/storage/blockvol/testrunner/scenarios/internal/recovery-baseline-failover.yaml --results-dir results/phase20-t6/stage1/recovery-baseline-failover ``` +Current reading: + +1. bootstrap closure inside `P20-T6-H1A` is now a prerequisite/setup step, not the pass/fail signal +2. the current red case is the sustained-workload path after `fio`, where large `dd_write` reproduces the default `64MB` WAL budget limitation +3. `Stage 1` should be rerun after WAL-size plumbing allows the master-managed create path to request a larger WAL budget + Preferred suite pack: ```bash @@ -394,9 +408,9 @@ It keeps semantic guardrails pinned while hardware work proceeds. ## Immediate Start Order 1. run `Stage 0` observation loop on the current baseline -2. start fixing the replica-membership wiring gap until `Stage 0` closes -3. run the `Stage 1` pack on existing YAMLs -4. only after `Stage 0` is closed and evidence transport is real, prepare Stage 2 YAML copies +2. keep the dedicated `P20-H0` bootstrap scenario green as a regression check +3. add WAL-size plumbing for the master-managed create path, then rerun the `Stage 1` pack +4. only after `Stage 1` has a valid WAL budget and evidence transport is real, prepare Stage 2 YAML copies ## What Not To Do diff --git a/sw-block/.private/phase/phase-20-test.md b/sw-block/.private/phase/phase-20-test.md index bbea53b5d..1accde1df 100644 --- a/sw-block/.private/phase/phase-20-test.md +++ b/sw-block/.private/phase/phase-20-test.md @@ -1,7 +1,7 @@ # Phase 20 Test Matrix Date: 2026-04-06 -Status: active; acceptance overlay added +Status: active; Stage 0 bootstrap closure split and frozen, Stage 1 workload follow-up pending ## Purpose @@ -256,7 +256,8 @@ Current tester automation status: 1. metadata-driven suite execution now exists for `Stage 0` and `Stage 1` 2. one command can now build, deploy, run remote scenarios, and collect evidence -3. the current `Stage 0` pipeline is operational but not yet passing end-to-end because the run still fails at the known first-write `dd_write` / `sync_all` barrier issue during `record-before` +3. `Stage 0` now has a clean bootstrap-only scenario and has passed on real hosts: `create -> 10s wait -> 4k fsync -> publish_healthy` +4. the remaining red hardware case is now isolated to `Stage 1`, where sustained `fio` plus large `dd_write` on the default `64MB` WAL budget reproduces the expected WAL-pressure workload failure Naming convention for this overlay: @@ -965,8 +966,8 @@ Execution details, scenario packs, and observation checklists live in: Goal: -1. fix the existing primary-core visibility gap so the promoted primary learns - its replica membership +1. prove the bootstrap membership gap is closed on real hosts +2. freeze the first-fence-to-`publish_healthy` chain as a standalone acceptance artifact Required hardware proof: @@ -987,8 +988,9 @@ Preferred suite entry: Current reading: 1. the suite pipeline itself now works end-to-end: build, deploy, remote execution, and evidence collection -2. the current `Stage 0` run is still red at `record-before` because of the known first-write `dd_write` / `sync_all` barrier issue -3. therefore `Stage 0` automation is real, but `Stage 0` acceptance is not yet closed +2. the clean `Stage 0` scenario now passes on real hosts: `create -> 10s wait -> 4k fsync -> wait_volume_healthy` +3. the observed bootstrap closure is `bootstrap-fence PASS` plus `wait_volume_healthy PASS after 5 polls (~10s)` which proves the `BarrierAccepted -> DurableLSN > 0 -> publish_healthy` chain +4. `Stage 0` should now be treated as closed bootstrap evidence, not as a carrier for sustained-workload failures ### Stage 1: V1 Failover + V2 Observation @@ -1020,6 +1022,12 @@ Preferred suite entry: 1. `sw-test-runner suite weed/storage/blockvol/testrunner/suites/phase20-t6-stage1.yaml` 2. `sw-test-runner suite weed/storage/blockvol/testrunner/suites/phase20-t6-stage1.yaml --skip-deploy` +Current reading: + +1. `Stage 1` is now intentionally separate from `Stage 0` and owns the sustained-I/O-under-established-replication question +2. the current failure is no longer bootstrap or protocol closure; it is the workload path after `fio` where the large `dd_write` hits `Input/output error` +3. the most likely current cause is default `64MB` WAL budget pressure on the master-managed create path under `sync_all`, so `Stage 1` should be rerun only after WAL-size plumbing exists + ### Stage 2: V2 Failover + V2 Decision Run with: @@ -1179,10 +1187,9 @@ Tier 1 component tests are implemented and passing. The current reading is: 3. **Tester acceptance overlay is now explicit.** `P20-A1..A6` define the minimum tester-owned cases needed to freeze the bounded product contract. -4. **Hardware automation is now real, but Stage 0 is still red.** The - metadata-driven suite can build, deploy, execute, and collect evidence, but - the current Stage 0 run still fails at the known first-write `dd_write` / - `sync_all` barrier issue. +4. **Hardware automation is now real, and Stage 0 is green.** The + metadata-driven suite can build, deploy, execute, and collect evidence, and + the bootstrap-only `Stage 0` scenario is now a passing real-host artifact. 5. **Hardware layer still follows the staged T6/T7 plan.** The automation path is better, but passing acceptance evidence still has to be earned on @@ -1191,7 +1198,7 @@ Tier 1 component tests are implemented and passing. The current reading is: Recommended next actions: 1. implement Tier 2 integration tests (3 tests, one new `qa_*` file) -2. close the known first-write `dd_write` / `sync_all` barrier blocker in `Stage 0` -3. rerun `sw-test-runner suite weed/storage/blockvol/testrunner/suites/phase20-t6-stage0.yaml` +2. add WAL size plumbing to the master-managed create path so `Stage 1` can request `256MB+` WAL +3. rerun `sw-test-runner suite weed/storage/blockvol/testrunner/suites/phase20-t6-stage1.yaml` 4. freeze tester-owned acceptance overlay `P20-A1..A6` 5. run the stage suites on hardware with V2 observation active diff --git a/sw-block/.private/phase/phase-20.md b/sw-block/.private/phase/phase-20.md index 5d640170f..0c46ea173 100644 --- a/sw-block/.private/phase/phase-20.md +++ b/sw-block/.private/phase/phase-20.md @@ -21,6 +21,70 @@ The V1 blockvol engine (`weed/storage/blockvol/`) — WAL, flusher, shipper, rebuild, iSCSI — stays untouched. Phase 20 changes who makes the *decision* (master failover logic, VS activation logic), not who *executes* it. +## Current Slice Replan + +Date: 2026-04-06 + +The current `Stage 1` hardware investigation changed one important assumption: +the remaining blocker is no longer "wire the existing V2 truth into the +production binary and leave `blockvol` semantics alone." + +The bounded bootstrap closure (`Stage 0`) is now closed. The active red case is +the post-bootstrap sustained-async-write path, where the current `CP13` runtime +still lets `1.5`-style local recovery autonomy interfere with the intended `v2` +control model. + +Current slice reading: + +1. `sync` / `fsync` / `SyncCache()` should be treated as control-plane durability fences, not as the data-plane replication mechanism +2. the replica may continue receiving and applying live-tail writes without any new durability confirmation being established +3. the current `CP13-6` retention max-bytes path uses `replicaFlushedLSN` (durability truth) as if it were the recoverability truth +4. this lets a local runtime budget transition the shipper into `needs_rebuild` before primary/replica negotiation has actually proven recoverability loss +5. that behavior matches a `1.5` local-autonomy assumption, not the intended `v2` ownership model + +Therefore the current slice plan is: + +1. remove `1.5` semantic ownership from local replica/shipper autonomy paths +2. preserve local execution machinery (`live shipping`, local flush, local catch-up executor, local rebuild executor) as host capabilities only +3. re-establish `catchup` and `rebuild` as negotiated `v2` control-plane outcomes between primary and replica +4. treat replica-local measurements as facts (`durable`, `received`, `applied`, `checkpoint`, `local pressure`, `local error`), not as final recovery decisions +5. require primary-visible negotiation before the system enters `catchup` or `needs_rebuild` as a semantic state + +This slice intentionally supersedes the earlier narrower assumption that +`Phase 20` would not need to change `blockvol` internals. The current blocker is +inside the recovery/control seam, so bounded `blockvol` changes are now in +scope when they are required to remove `1.5` semantic ownership and restore the +intended `v2` negotiated model. + +### Current Slice Goals + +1. separate durability truth from recoverability truth +2. stop using local retention-budget heuristics as autonomous rebuild authority +3. make `sync` the place where the primary learns whether the replica is in `keepup`, `catchup`, or `rebuild_required` +4. ensure `needs_rebuild` is reached only after recoverability loss is proven, not just inferred from missing barrier progress during async writes +5. keep local pressure protection and fail-closed durability semantics intact while moving recovery-state ownership back to negotiated `v2` control flow + +### Current Slice Non-Goals + +1. do not redefine `WriteLBA()` as a durability API +2. do not remove local flush, WAL pressure handling, or other host protection mechanisms +3. do not silently relax `sync_all` durability guarantees +4. do not let replica-local heuristics directly set outward semantic truth +5. do not broaden this slice into a full new transport or a broad rebuild redesign before the control ownership is corrected + +### Current Slice Exit Criteria + +This slice is complete only when: + +1. local replica/shipper code reports facts and bounded hints, but no longer unilaterally owns semantic `catchup` / `needs_rebuild` transitions +2. primary/replica recovery progression is explicit enough that `catchup` can be entered and pinned without immediately collapsing into rebuild from a local budget threshold alone +3. `needs_rebuild` is reached only after negotiated evidence shows the recoverable envelope is actually lost +4. `Stage 1` failures can be read as either: + - real recoverability loss proved by negotiated evidence, or + - bounded `catchup` not yet sufficient, + but not as a local-autonomy side effect hidden behind `1.5` logic +5. the phase log carries the concrete technical design and implementation slice boundaries for this replan + ## V2 Promise (Non-Negotiable) These constraints govern every task in this phase and every future phase. @@ -113,7 +177,7 @@ name. Code and API must make the distinction explicit. 1. Every task must change code in `weed/server/` or `weed/storage/blockvol/` 2. No new files in `sw-block/runtime/volumev2/` unless adapter stubs 3. Validation is `sw-test-runner` on m01/M02, not new POC tests -4. The V1 blockvol engine must not change and must not regress +4. Broad `blockvol` behavior must not regress; only bounded ownership/recovery-seam changes are allowed in the current slice replan 5. Fresh promotion evidence is mandatory for V2-mode failover 6. Durability-first candidate selection is mandatory 7. Local activation gate is mandatory before serving @@ -364,7 +428,7 @@ T1 and T2 can run in parallel. T3, T4, T5 can partially overlap. 1. Replace heartbeat gRPC transport — it stays 2. Replace WAL shipper — it stays (V1 execution, V2 orchestrated) 3. Replace assignment queue — it stays -4. Change blockvol engine internals — WAL/flusher/rebuild untouched +4. Broaden changes beyond the bounded ownership/recovery seam now required by the current slice replan 5. Build new simulation tests — sw-test-runner is the oracle 6. Add new files to `sw-block/runtime/volumev2/` diff --git a/weed/pb/master.proto b/weed/pb/master.proto index 3175b894c..3558fca37 100644 --- a/weed/pb/master.proto +++ b/weed/pb/master.proto @@ -554,6 +554,7 @@ message CreateBlockVolumeRequest { string disk_type = 3; uint32 replica_factor = 4; string durability_mode = 5; + uint64 wal_size_bytes = 6; } message CreateBlockVolumeResponse { string volume_id = 1; diff --git a/weed/pb/master_pb/master.pb.go b/weed/pb/master_pb/master.pb.go index 62028cdef..f8d138c6f 100644 --- a/weed/pb/master_pb/master.pb.go +++ b/weed/pb/master_pb/master.pb.go @@ -1781,7 +1781,9 @@ func (x *StatisticsResponse) GetFileCount() uint64 { return 0 } +// // collection related +// type Collection struct { state protoimpl.MessageState `protogen:"open.v1"` Name string `protobuf:"bytes,1,opt,name=name,proto3" json:"name,omitempty"` @@ -2002,7 +2004,9 @@ func (*CollectionDeleteResponse) Descriptor() ([]byte, []int) { return file_master_proto_rawDescGZIP(), []int{24} } +// // volume related +// type DiskInfo struct { state protoimpl.MessageState `protogen:"open.v1"` Type string `protobuf:"bytes,1,opt,name=type,proto3" json:"type,omitempty"` @@ -3900,35 +3904,35 @@ func (*VolumeGrowResponse) Descriptor() ([]byte, []int) { } type BlockVolumeInfoMessage struct { - state protoimpl.MessageState `protogen:"open.v1"` - Path string `protobuf:"bytes,1,opt,name=path,proto3" json:"path,omitempty"` - VolumeSize uint64 `protobuf:"varint,2,opt,name=volume_size,json=volumeSize,proto3" json:"volume_size,omitempty"` - BlockSize uint32 `protobuf:"varint,3,opt,name=block_size,json=blockSize,proto3" json:"block_size,omitempty"` - Epoch uint64 `protobuf:"varint,4,opt,name=epoch,proto3" json:"epoch,omitempty"` - Role uint32 `protobuf:"varint,5,opt,name=role,proto3" json:"role,omitempty"` - WalHeadLsn uint64 `protobuf:"varint,6,opt,name=wal_head_lsn,json=walHeadLsn,proto3" json:"wal_head_lsn,omitempty"` - CheckpointLsn uint64 `protobuf:"varint,7,opt,name=checkpoint_lsn,json=checkpointLsn,proto3" json:"checkpoint_lsn,omitempty"` - HasLease bool `protobuf:"varint,8,opt,name=has_lease,json=hasLease,proto3" json:"has_lease,omitempty"` - DiskType string `protobuf:"bytes,9,opt,name=disk_type,json=diskType,proto3" json:"disk_type,omitempty"` - ReplicaDataAddr string `protobuf:"bytes,10,opt,name=replica_data_addr,json=replicaDataAddr,proto3" json:"replica_data_addr,omitempty"` - ReplicaCtrlAddr string `protobuf:"bytes,11,opt,name=replica_ctrl_addr,json=replicaCtrlAddr,proto3" json:"replica_ctrl_addr,omitempty"` - HealthScore float64 `protobuf:"fixed64,12,opt,name=health_score,json=healthScore,proto3" json:"health_score,omitempty"` - ScrubErrors int64 `protobuf:"varint,13,opt,name=scrub_errors,json=scrubErrors,proto3" json:"scrub_errors,omitempty"` - LastScrubTime int64 `protobuf:"varint,14,opt,name=last_scrub_time,json=lastScrubTime,proto3" json:"last_scrub_time,omitempty"` - ReplicaDegraded bool `protobuf:"varint,15,opt,name=replica_degraded,json=replicaDegraded,proto3" json:"replica_degraded,omitempty"` - DurabilityMode string `protobuf:"bytes,16,opt,name=durability_mode,json=durabilityMode,proto3" json:"durability_mode,omitempty"` - NvmeAddr string `protobuf:"bytes,17,opt,name=nvme_addr,json=nvmeAddr,proto3" json:"nvme_addr,omitempty"` - Nqn string `protobuf:"bytes,18,opt,name=nqn,proto3" json:"nqn,omitempty"` - ReplicaReady *bool `protobuf:"varint,19,opt,name=replica_ready,json=replicaReady,proto3,oneof" json:"replica_ready,omitempty"` - NeedsRebuild *bool `protobuf:"varint,20,opt,name=needs_rebuild,json=needsRebuild,proto3,oneof" json:"needs_rebuild,omitempty"` - PublishHealthy *bool `protobuf:"varint,21,opt,name=publish_healthy,json=publishHealthy,proto3,oneof" json:"publish_healthy,omitempty"` - VolumeMode *string `protobuf:"bytes,22,opt,name=volume_mode,json=volumeMode,proto3,oneof" json:"volume_mode,omitempty"` - VolumeModeReason *string `protobuf:"bytes,23,opt,name=volume_mode_reason,json=volumeModeReason,proto3,oneof" json:"volume_mode_reason,omitempty"` - EngineProjectionMode *string `protobuf:"bytes,24,opt,name=engine_projection_mode,json=engineProjectionMode,proto3,oneof" json:"engine_projection_mode,omitempty"` - ActivationGated *bool `protobuf:"varint,25,opt,name=activation_gated,json=activationGated,proto3,oneof" json:"activation_gated,omitempty"` - ActivationGateReason *string `protobuf:"bytes,26,opt,name=activation_gate_reason,json=activationGateReason,proto3,oneof" json:"activation_gate_reason,omitempty"` - unknownFields protoimpl.UnknownFields - sizeCache protoimpl.SizeCache + state protoimpl.MessageState `protogen:"open.v1"` + Path string `protobuf:"bytes,1,opt,name=path,proto3" json:"path,omitempty"` + VolumeSize uint64 `protobuf:"varint,2,opt,name=volume_size,json=volumeSize,proto3" json:"volume_size,omitempty"` + BlockSize uint32 `protobuf:"varint,3,opt,name=block_size,json=blockSize,proto3" json:"block_size,omitempty"` + Epoch uint64 `protobuf:"varint,4,opt,name=epoch,proto3" json:"epoch,omitempty"` + Role uint32 `protobuf:"varint,5,opt,name=role,proto3" json:"role,omitempty"` + WalHeadLsn uint64 `protobuf:"varint,6,opt,name=wal_head_lsn,json=walHeadLsn,proto3" json:"wal_head_lsn,omitempty"` + CheckpointLsn uint64 `protobuf:"varint,7,opt,name=checkpoint_lsn,json=checkpointLsn,proto3" json:"checkpoint_lsn,omitempty"` + HasLease bool `protobuf:"varint,8,opt,name=has_lease,json=hasLease,proto3" json:"has_lease,omitempty"` + DiskType string `protobuf:"bytes,9,opt,name=disk_type,json=diskType,proto3" json:"disk_type,omitempty"` + ReplicaDataAddr string `protobuf:"bytes,10,opt,name=replica_data_addr,json=replicaDataAddr,proto3" json:"replica_data_addr,omitempty"` + ReplicaCtrlAddr string `protobuf:"bytes,11,opt,name=replica_ctrl_addr,json=replicaCtrlAddr,proto3" json:"replica_ctrl_addr,omitempty"` + HealthScore float64 `protobuf:"fixed64,12,opt,name=health_score,json=healthScore,proto3" json:"health_score,omitempty"` + ScrubErrors int64 `protobuf:"varint,13,opt,name=scrub_errors,json=scrubErrors,proto3" json:"scrub_errors,omitempty"` + LastScrubTime int64 `protobuf:"varint,14,opt,name=last_scrub_time,json=lastScrubTime,proto3" json:"last_scrub_time,omitempty"` + ReplicaDegraded bool `protobuf:"varint,15,opt,name=replica_degraded,json=replicaDegraded,proto3" json:"replica_degraded,omitempty"` + DurabilityMode string `protobuf:"bytes,16,opt,name=durability_mode,json=durabilityMode,proto3" json:"durability_mode,omitempty"` + NvmeAddr string `protobuf:"bytes,17,opt,name=nvme_addr,json=nvmeAddr,proto3" json:"nvme_addr,omitempty"` + Nqn string `protobuf:"bytes,18,opt,name=nqn,proto3" json:"nqn,omitempty"` + ReplicaReady *bool `protobuf:"varint,19,opt,name=replica_ready,json=replicaReady,proto3,oneof" json:"replica_ready,omitempty"` + NeedsRebuild *bool `protobuf:"varint,20,opt,name=needs_rebuild,json=needsRebuild,proto3,oneof" json:"needs_rebuild,omitempty"` + PublishHealthy *bool `protobuf:"varint,21,opt,name=publish_healthy,json=publishHealthy,proto3,oneof" json:"publish_healthy,omitempty"` + VolumeMode *string `protobuf:"bytes,22,opt,name=volume_mode,json=volumeMode,proto3,oneof" json:"volume_mode,omitempty"` + VolumeModeReason *string `protobuf:"bytes,23,opt,name=volume_mode_reason,json=volumeModeReason,proto3,oneof" json:"volume_mode_reason,omitempty"` + EngineProjectionMode *string `protobuf:"bytes,24,opt,name=engine_projection_mode,json=engineProjectionMode,proto3,oneof" json:"engine_projection_mode,omitempty"` // V2: pure engine-derived local projection mode + ActivationGated *bool `protobuf:"varint,25,opt,name=activation_gated,json=activationGated,proto3,oneof" json:"activation_gated,omitempty"` // T4: true if activation-gated from serving + ActivationGateReason *string `protobuf:"bytes,26,opt,name=activation_gate_reason,json=activationGateReason,proto3,oneof" json:"activation_gate_reason,omitempty"` // T4: reason for activation gate + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache } func (x *BlockVolumeInfoMessage) Reset() { @@ -4386,6 +4390,7 @@ type CreateBlockVolumeRequest struct { DiskType string `protobuf:"bytes,3,opt,name=disk_type,json=diskType,proto3" json:"disk_type,omitempty"` ReplicaFactor uint32 `protobuf:"varint,4,opt,name=replica_factor,json=replicaFactor,proto3" json:"replica_factor,omitempty"` DurabilityMode string `protobuf:"bytes,5,opt,name=durability_mode,json=durabilityMode,proto3" json:"durability_mode,omitempty"` + WalSizeBytes uint64 `protobuf:"varint,6,opt,name=wal_size_bytes,json=walSizeBytes,proto3" json:"wal_size_bytes,omitempty"` unknownFields protoimpl.UnknownFields sizeCache protoimpl.SizeCache } @@ -4455,6 +4460,13 @@ func (x *CreateBlockVolumeRequest) GetDurabilityMode() string { return "" } +func (x *CreateBlockVolumeRequest) GetWalSizeBytes() uint64 { + if x != nil { + return x.WalSizeBytes + } + return 0 +} + type CreateBlockVolumeResponse struct { state protoimpl.MessageState `protogen:"open.v1"` VolumeId string `protobuf:"bytes,1,opt,name=volume_id,json=volumeId,proto3" json:"volume_id,omitempty"` @@ -6032,7 +6044,7 @@ const file_master_proto_rawDesc = "" + "\x0fprevious_leader\x18\x01 \x01(\tR\x0epreviousLeader\x12\x1d\n" + "\n" + "new_leader\x18\x02 \x01(\tR\tnewLeader\"\x14\n" + - "\x12VolumeGrowResponse\"\x9c\a\n" + + "\x12VolumeGrowResponse\"\x8d\t\n" + "\x16BlockVolumeInfoMessage\x12\x12\n" + "\x04path\x18\x01 \x01(\tR\x04path\x12\x1f\n" + "\vvolume_size\x18\x02 \x01(\x04R\n" + @@ -6061,12 +6073,18 @@ const file_master_proto_rawDesc = "" + "\x0fpublish_healthy\x18\x15 \x01(\bH\x02R\x0epublishHealthy\x88\x01\x01\x12$\n" + "\vvolume_mode\x18\x16 \x01(\tH\x03R\n" + "volumeMode\x88\x01\x01\x121\n" + - "\x12volume_mode_reason\x18\x17 \x01(\tH\x04R\x10volumeModeReason\x88\x01\x01B\x10\n" + + "\x12volume_mode_reason\x18\x17 \x01(\tH\x04R\x10volumeModeReason\x88\x01\x01\x129\n" + + "\x16engine_projection_mode\x18\x18 \x01(\tH\x05R\x14engineProjectionMode\x88\x01\x01\x12.\n" + + "\x10activation_gated\x18\x19 \x01(\bH\x06R\x0factivationGated\x88\x01\x01\x129\n" + + "\x16activation_gate_reason\x18\x1a \x01(\tH\aR\x14activationGateReason\x88\x01\x01B\x10\n" + "\x0e_replica_readyB\x10\n" + "\x0e_needs_rebuildB\x12\n" + "\x10_publish_healthyB\x0e\n" + "\f_volume_modeB\x15\n" + - "\x13_volume_mode_reason\"\x8e\x01\n" + + "\x13_volume_mode_reasonB\x19\n" + + "\x17_engine_projection_modeB\x13\n" + + "\x11_activation_gatedB\x19\n" + + "\x17_activation_gate_reason\"\x8e\x01\n" + "\x1bBlockVolumeShortInfoMessage\x12\x12\n" + "\x04path\x18\x01 \x01(\tR\x04path\x12\x1f\n" + "\vvolume_size\x18\x02 \x01(\x04R\n" + @@ -6088,14 +6106,15 @@ const file_master_proto_rawDesc = "" + "\x12ReplicaAddrMessage\x12\x1b\n" + "\tdata_addr\x18\x01 \x01(\tR\bdataAddr\x12\x1b\n" + "\tctrl_addr\x18\x02 \x01(\tR\bctrlAddr\x12\x1b\n" + - "\tserver_id\x18\x03 \x01(\tR\bserverId\"\xba\x01\n" + + "\tserver_id\x18\x03 \x01(\tR\bserverId\"\xe0\x01\n" + "\x18CreateBlockVolumeRequest\x12\x12\n" + "\x04name\x18\x01 \x01(\tR\x04name\x12\x1d\n" + "\n" + "size_bytes\x18\x02 \x01(\x04R\tsizeBytes\x12\x1b\n" + "\tdisk_type\x18\x03 \x01(\tR\bdiskType\x12%\n" + "\x0ereplica_factor\x18\x04 \x01(\rR\rreplicaFactor\x12'\n" + - "\x0fdurability_mode\x18\x05 \x01(\tR\x0edurabilityMode\"\xb4\x02\n" + + "\x0fdurability_mode\x18\x05 \x01(\tR\x0edurabilityMode\x12$\n" + + "\x0ewal_size_bytes\x18\x06 \x01(\x04R\fwalSizeBytes\"\xb4\x02\n" + "\x19CreateBlockVolumeResponse\x12\x1b\n" + "\tvolume_id\x18\x01 \x01(\tR\bvolumeId\x12#\n" + "\rvolume_server\x18\x02 \x01(\tR\fvolumeServer\x12\x1d\n" + diff --git a/weed/pb/volume_server.proto b/weed/pb/volume_server.proto index d18055199..6cbbc38ef 100644 --- a/weed/pb/volume_server.proto +++ b/weed/pb/volume_server.proto @@ -790,6 +790,7 @@ message AllocateBlockVolumeRequest { uint64 size_bytes = 2; string disk_type = 3; string durability_mode = 4; + uint64 wal_size_bytes = 5; } message AllocateBlockVolumeResponse { string path = 1; diff --git a/weed/pb/volume_server_pb/volume_server.pb.go b/weed/pb/volume_server_pb/volume_server.pb.go index 124229436..1609f7eb9 100644 --- a/weed/pb/volume_server_pb/volume_server.pb.go +++ b/weed/pb/volume_server_pb/volume_server.pb.go @@ -6191,6 +6191,7 @@ type AllocateBlockVolumeRequest struct { SizeBytes uint64 `protobuf:"varint,2,opt,name=size_bytes,json=sizeBytes,proto3" json:"size_bytes,omitempty"` DiskType string `protobuf:"bytes,3,opt,name=disk_type,json=diskType,proto3" json:"disk_type,omitempty"` DurabilityMode string `protobuf:"bytes,4,opt,name=durability_mode,json=durabilityMode,proto3" json:"durability_mode,omitempty"` + WalSizeBytes uint64 `protobuf:"varint,5,opt,name=wal_size_bytes,json=walSizeBytes,proto3" json:"wal_size_bytes,omitempty"` unknownFields protoimpl.UnknownFields sizeCache protoimpl.SizeCache } @@ -6253,6 +6254,13 @@ func (x *AllocateBlockVolumeRequest) GetDurabilityMode() string { return "" } +func (x *AllocateBlockVolumeRequest) GetWalSizeBytes() uint64 { + if x != nil { + return x.WalSizeBytes + } + return 0 +} + type AllocateBlockVolumeResponse struct { state protoimpl.MessageState `protogen:"open.v1"` Path string `protobuf:"bytes,1,opt,name=path,proto3" json:"path,omitempty"` @@ -7245,6 +7253,168 @@ func (*CancelExpandBlockVolumeResponse) Descriptor() ([]byte, []int) { return file_volume_server_proto_rawDescGZIP(), []int{127} } +// T2: Fresh on-demand promotion evidence. Master queries VS at failover +// time. Returns live local facts, not cached/stale heartbeat data. +type QueryBlockPromotionEvidenceRequest struct { + state protoimpl.MessageState `protogen:"open.v1"` + Path string `protobuf:"bytes,1,opt,name=path,proto3" json:"path,omitempty"` // volume file path on queried VS + ExpectedEpoch uint64 `protobuf:"varint,2,opt,name=expected_epoch,json=expectedEpoch,proto3" json:"expected_epoch,omitempty"` // caller's expected epoch for staleness check + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache +} + +func (x *QueryBlockPromotionEvidenceRequest) Reset() { + *x = QueryBlockPromotionEvidenceRequest{} + mi := &file_volume_server_proto_msgTypes[128] + ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) + ms.StoreMessageInfo(mi) +} + +func (x *QueryBlockPromotionEvidenceRequest) String() string { + return protoimpl.X.MessageStringOf(x) +} + +func (*QueryBlockPromotionEvidenceRequest) ProtoMessage() {} + +func (x *QueryBlockPromotionEvidenceRequest) ProtoReflect() protoreflect.Message { + mi := &file_volume_server_proto_msgTypes[128] + if x != nil { + ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) + if ms.LoadMessageInfo() == nil { + ms.StoreMessageInfo(mi) + } + return ms + } + return mi.MessageOf(x) +} + +// Deprecated: Use QueryBlockPromotionEvidenceRequest.ProtoReflect.Descriptor instead. +func (*QueryBlockPromotionEvidenceRequest) Descriptor() ([]byte, []int) { + return file_volume_server_proto_rawDescGZIP(), []int{128} +} + +func (x *QueryBlockPromotionEvidenceRequest) GetPath() string { + if x != nil { + return x.Path + } + return "" +} + +func (x *QueryBlockPromotionEvidenceRequest) GetExpectedEpoch() uint64 { + if x != nil { + return x.ExpectedEpoch + } + return 0 +} + +type QueryBlockPromotionEvidenceResponse struct { + state protoimpl.MessageState `protogen:"open.v1"` + Path string `protobuf:"bytes,1,opt,name=path,proto3" json:"path,omitempty"` + Epoch uint64 `protobuf:"varint,2,opt,name=epoch,proto3" json:"epoch,omitempty"` + CommittedLsn uint64 `protobuf:"varint,3,opt,name=committed_lsn,json=committedLsn,proto3" json:"committed_lsn,omitempty"` + WalHeadLsn uint64 `protobuf:"varint,4,opt,name=wal_head_lsn,json=walHeadLsn,proto3" json:"wal_head_lsn,omitempty"` + CheckpointLsn uint64 `protobuf:"varint,5,opt,name=checkpoint_lsn,json=checkpointLsn,proto3" json:"checkpoint_lsn,omitempty"` + EngineProjectionMode string `protobuf:"bytes,6,opt,name=engine_projection_mode,json=engineProjectionMode,proto3" json:"engine_projection_mode,omitempty"` // pure V2 engine local projection, empty if no core + Eligible bool `protobuf:"varint,7,opt,name=eligible,proto3" json:"eligible,omitempty"` + Reason string `protobuf:"bytes,8,opt,name=reason,proto3" json:"reason,omitempty"` // why ineligible, or empty + HealthScore float64 `protobuf:"fixed64,9,opt,name=health_score,json=healthScore,proto3" json:"health_score,omitempty"` + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache +} + +func (x *QueryBlockPromotionEvidenceResponse) Reset() { + *x = QueryBlockPromotionEvidenceResponse{} + mi := &file_volume_server_proto_msgTypes[129] + ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) + ms.StoreMessageInfo(mi) +} + +func (x *QueryBlockPromotionEvidenceResponse) String() string { + return protoimpl.X.MessageStringOf(x) +} + +func (*QueryBlockPromotionEvidenceResponse) ProtoMessage() {} + +func (x *QueryBlockPromotionEvidenceResponse) ProtoReflect() protoreflect.Message { + mi := &file_volume_server_proto_msgTypes[129] + if x != nil { + ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) + if ms.LoadMessageInfo() == nil { + ms.StoreMessageInfo(mi) + } + return ms + } + return mi.MessageOf(x) +} + +// Deprecated: Use QueryBlockPromotionEvidenceResponse.ProtoReflect.Descriptor instead. +func (*QueryBlockPromotionEvidenceResponse) Descriptor() ([]byte, []int) { + return file_volume_server_proto_rawDescGZIP(), []int{129} +} + +func (x *QueryBlockPromotionEvidenceResponse) GetPath() string { + if x != nil { + return x.Path + } + return "" +} + +func (x *QueryBlockPromotionEvidenceResponse) GetEpoch() uint64 { + if x != nil { + return x.Epoch + } + return 0 +} + +func (x *QueryBlockPromotionEvidenceResponse) GetCommittedLsn() uint64 { + if x != nil { + return x.CommittedLsn + } + return 0 +} + +func (x *QueryBlockPromotionEvidenceResponse) GetWalHeadLsn() uint64 { + if x != nil { + return x.WalHeadLsn + } + return 0 +} + +func (x *QueryBlockPromotionEvidenceResponse) GetCheckpointLsn() uint64 { + if x != nil { + return x.CheckpointLsn + } + return 0 +} + +func (x *QueryBlockPromotionEvidenceResponse) GetEngineProjectionMode() string { + if x != nil { + return x.EngineProjectionMode + } + return "" +} + +func (x *QueryBlockPromotionEvidenceResponse) GetEligible() bool { + if x != nil { + return x.Eligible + } + return false +} + +func (x *QueryBlockPromotionEvidenceResponse) GetReason() string { + if x != nil { + return x.Reason + } + return "" +} + +func (x *QueryBlockPromotionEvidenceResponse) GetHealthScore() float64 { + if x != nil { + return x.HealthScore + } + return 0 +} + type FetchAndWriteNeedleRequest_Replica struct { state protoimpl.MessageState `protogen:"open.v1"` Url string `protobuf:"bytes,1,opt,name=url,proto3" json:"url,omitempty"` @@ -7256,7 +7426,7 @@ type FetchAndWriteNeedleRequest_Replica struct { func (x *FetchAndWriteNeedleRequest_Replica) Reset() { *x = FetchAndWriteNeedleRequest_Replica{} - mi := &file_volume_server_proto_msgTypes[128] + mi := &file_volume_server_proto_msgTypes[130] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -7268,7 +7438,7 @@ func (x *FetchAndWriteNeedleRequest_Replica) String() string { func (*FetchAndWriteNeedleRequest_Replica) ProtoMessage() {} func (x *FetchAndWriteNeedleRequest_Replica) ProtoReflect() protoreflect.Message { - mi := &file_volume_server_proto_msgTypes[128] + mi := &file_volume_server_proto_msgTypes[130] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -7316,7 +7486,7 @@ type QueryRequest_Filter struct { func (x *QueryRequest_Filter) Reset() { *x = QueryRequest_Filter{} - mi := &file_volume_server_proto_msgTypes[129] + mi := &file_volume_server_proto_msgTypes[131] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -7328,7 +7498,7 @@ func (x *QueryRequest_Filter) String() string { func (*QueryRequest_Filter) ProtoMessage() {} func (x *QueryRequest_Filter) ProtoReflect() protoreflect.Message { - mi := &file_volume_server_proto_msgTypes[129] + mi := &file_volume_server_proto_msgTypes[131] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -7378,7 +7548,7 @@ type QueryRequest_InputSerialization struct { func (x *QueryRequest_InputSerialization) Reset() { *x = QueryRequest_InputSerialization{} - mi := &file_volume_server_proto_msgTypes[130] + mi := &file_volume_server_proto_msgTypes[132] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -7390,7 +7560,7 @@ func (x *QueryRequest_InputSerialization) String() string { func (*QueryRequest_InputSerialization) ProtoMessage() {} func (x *QueryRequest_InputSerialization) ProtoReflect() protoreflect.Message { - mi := &file_volume_server_proto_msgTypes[130] + mi := &file_volume_server_proto_msgTypes[132] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -7444,7 +7614,7 @@ type QueryRequest_OutputSerialization struct { func (x *QueryRequest_OutputSerialization) Reset() { *x = QueryRequest_OutputSerialization{} - mi := &file_volume_server_proto_msgTypes[131] + mi := &file_volume_server_proto_msgTypes[133] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -7456,7 +7626,7 @@ func (x *QueryRequest_OutputSerialization) String() string { func (*QueryRequest_OutputSerialization) ProtoMessage() {} func (x *QueryRequest_OutputSerialization) ProtoReflect() protoreflect.Message { - mi := &file_volume_server_proto_msgTypes[131] + mi := &file_volume_server_proto_msgTypes[133] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -7502,7 +7672,7 @@ type QueryRequest_InputSerialization_CSVInput struct { func (x *QueryRequest_InputSerialization_CSVInput) Reset() { *x = QueryRequest_InputSerialization_CSVInput{} - mi := &file_volume_server_proto_msgTypes[132] + mi := &file_volume_server_proto_msgTypes[134] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -7514,7 +7684,7 @@ func (x *QueryRequest_InputSerialization_CSVInput) String() string { func (*QueryRequest_InputSerialization_CSVInput) ProtoMessage() {} func (x *QueryRequest_InputSerialization_CSVInput) ProtoReflect() protoreflect.Message { - mi := &file_volume_server_proto_msgTypes[132] + mi := &file_volume_server_proto_msgTypes[134] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -7588,7 +7758,7 @@ type QueryRequest_InputSerialization_JSONInput struct { func (x *QueryRequest_InputSerialization_JSONInput) Reset() { *x = QueryRequest_InputSerialization_JSONInput{} - mi := &file_volume_server_proto_msgTypes[133] + mi := &file_volume_server_proto_msgTypes[135] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -7600,7 +7770,7 @@ func (x *QueryRequest_InputSerialization_JSONInput) String() string { func (*QueryRequest_InputSerialization_JSONInput) ProtoMessage() {} func (x *QueryRequest_InputSerialization_JSONInput) ProtoReflect() protoreflect.Message { - mi := &file_volume_server_proto_msgTypes[133] + mi := &file_volume_server_proto_msgTypes[135] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -7631,7 +7801,7 @@ type QueryRequest_InputSerialization_ParquetInput struct { func (x *QueryRequest_InputSerialization_ParquetInput) Reset() { *x = QueryRequest_InputSerialization_ParquetInput{} - mi := &file_volume_server_proto_msgTypes[134] + mi := &file_volume_server_proto_msgTypes[136] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -7643,7 +7813,7 @@ func (x *QueryRequest_InputSerialization_ParquetInput) String() string { func (*QueryRequest_InputSerialization_ParquetInput) ProtoMessage() {} func (x *QueryRequest_InputSerialization_ParquetInput) ProtoReflect() protoreflect.Message { - mi := &file_volume_server_proto_msgTypes[134] + mi := &file_volume_server_proto_msgTypes[136] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -7672,7 +7842,7 @@ type QueryRequest_OutputSerialization_CSVOutput struct { func (x *QueryRequest_OutputSerialization_CSVOutput) Reset() { *x = QueryRequest_OutputSerialization_CSVOutput{} - mi := &file_volume_server_proto_msgTypes[135] + mi := &file_volume_server_proto_msgTypes[137] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -7684,7 +7854,7 @@ func (x *QueryRequest_OutputSerialization_CSVOutput) String() string { func (*QueryRequest_OutputSerialization_CSVOutput) ProtoMessage() {} func (x *QueryRequest_OutputSerialization_CSVOutput) ProtoReflect() protoreflect.Message { - mi := &file_volume_server_proto_msgTypes[135] + mi := &file_volume_server_proto_msgTypes[137] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -7744,7 +7914,7 @@ type QueryRequest_OutputSerialization_JSONOutput struct { func (x *QueryRequest_OutputSerialization_JSONOutput) Reset() { *x = QueryRequest_OutputSerialization_JSONOutput{} - mi := &file_volume_server_proto_msgTypes[136] + mi := &file_volume_server_proto_msgTypes[138] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -7756,7 +7926,7 @@ func (x *QueryRequest_OutputSerialization_JSONOutput) String() string { func (*QueryRequest_OutputSerialization_JSONOutput) ProtoMessage() {} func (x *QueryRequest_OutputSerialization_JSONOutput) ProtoReflect() protoreflect.Message { - mi := &file_volume_server_proto_msgTypes[136] + mi := &file_volume_server_proto_msgTypes[138] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -8282,13 +8452,14 @@ const file_volume_server_proto_rawDesc = "" + "\rstart_time_ns\x18\x01 \x01(\x03R\vstartTimeNs\x12$\n" + "\x0eremote_time_ns\x18\x02 \x01(\x03R\fremoteTimeNs\x12 \n" + "\fstop_time_ns\x18\x03 \x01(\x03R\n" + - "stopTimeNs\"\x95\x01\n" + + "stopTimeNs\"\xbb\x01\n" + "\x1aAllocateBlockVolumeRequest\x12\x12\n" + "\x04name\x18\x01 \x01(\tR\x04name\x12\x1d\n" + "\n" + "size_bytes\x18\x02 \x01(\x04R\tsizeBytes\x12\x1b\n" + "\tdisk_type\x18\x03 \x01(\tR\bdiskType\x12'\n" + - "\x0fdurability_mode\x18\x04 \x01(\tR\x0edurabilityMode\"\x99\x02\n" + + "\x0fdurability_mode\x18\x04 \x01(\tR\x0edurabilityMode\x12$\n" + + "\x0ewal_size_bytes\x18\x05 \x01(\x04R\fwalSizeBytes\"\x99\x02\n" + "\x1bAllocateBlockVolumeResponse\x12\x12\n" + "\x04path\x18\x01 \x01(\tR\x04path\x12\x10\n" + "\x03iqn\x18\x02 \x01(\tR\x03iqn\x12\x1d\n" + @@ -8351,12 +8522,26 @@ const file_volume_server_proto_rawDesc = "" + "\x1eCancelExpandBlockVolumeRequest\x12\x12\n" + "\x04name\x18\x01 \x01(\tR\x04name\x12!\n" + "\fexpand_epoch\x18\x02 \x01(\x04R\vexpandEpoch\"!\n" + - "\x1fCancelExpandBlockVolumeResponse*>\n" + + "\x1fCancelExpandBlockVolumeResponse\"_\n" + + "\"QueryBlockPromotionEvidenceRequest\x12\x12\n" + + "\x04path\x18\x01 \x01(\tR\x04path\x12%\n" + + "\x0eexpected_epoch\x18\x02 \x01(\x04R\rexpectedEpoch\"\xca\x02\n" + + "#QueryBlockPromotionEvidenceResponse\x12\x12\n" + + "\x04path\x18\x01 \x01(\tR\x04path\x12\x14\n" + + "\x05epoch\x18\x02 \x01(\x04R\x05epoch\x12#\n" + + "\rcommitted_lsn\x18\x03 \x01(\x04R\fcommittedLsn\x12 \n" + + "\fwal_head_lsn\x18\x04 \x01(\x04R\n" + + "walHeadLsn\x12%\n" + + "\x0echeckpoint_lsn\x18\x05 \x01(\x04R\rcheckpointLsn\x124\n" + + "\x16engine_projection_mode\x18\x06 \x01(\tR\x14engineProjectionMode\x12\x1a\n" + + "\beligible\x18\a \x01(\bR\beligible\x12\x16\n" + + "\x06reason\x18\b \x01(\tR\x06reason\x12!\n" + + "\fhealth_score\x18\t \x01(\x01R\vhealthScore*>\n" + "\x0fVolumeScrubMode\x12\v\n" + "\aUNKNOWN\x10\x00\x12\t\n" + "\x05INDEX\x10\x01\x12\b\n" + "\x04FULL\x10\x02\x12\t\n" + - "\x05LOCAL\x10\x032\xda2\n" + + "\x05LOCAL\x10\x032\xe93\n" + "\fVolumeServer\x12\\\n" + "\vBatchDelete\x12$.volume_server_pb.BatchDeleteRequest\x1a%.volume_server_pb.BatchDeleteResponse\"\x00\x12n\n" + "\x11VacuumVolumeCheck\x12*.volume_server_pb.VacuumVolumeCheckRequest\x1a+.volume_server_pb.VacuumVolumeCheckResponse\"\x00\x12v\n" + @@ -8416,7 +8601,8 @@ const file_volume_server_proto_rawDesc = "" + "\x11ExpandBlockVolume\x12*.volume_server_pb.ExpandBlockVolumeRequest\x1a+.volume_server_pb.ExpandBlockVolumeResponse\"\x00\x12\x83\x01\n" + "\x18PrepareExpandBlockVolume\x121.volume_server_pb.PrepareExpandBlockVolumeRequest\x1a2.volume_server_pb.PrepareExpandBlockVolumeResponse\"\x00\x12\x80\x01\n" + "\x17CommitExpandBlockVolume\x120.volume_server_pb.CommitExpandBlockVolumeRequest\x1a1.volume_server_pb.CommitExpandBlockVolumeResponse\"\x00\x12\x80\x01\n" + - "\x17CancelExpandBlockVolume\x120.volume_server_pb.CancelExpandBlockVolumeRequest\x1a1.volume_server_pb.CancelExpandBlockVolumeResponse\"\x00B9Z7github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pbb\x06proto3" + "\x17CancelExpandBlockVolume\x120.volume_server_pb.CancelExpandBlockVolumeRequest\x1a1.volume_server_pb.CancelExpandBlockVolumeResponse\"\x00\x12\x8c\x01\n" + + "\x1bQueryBlockPromotionEvidence\x124.volume_server_pb.QueryBlockPromotionEvidenceRequest\x1a5.volume_server_pb.QueryBlockPromotionEvidenceResponse\"\x00B9Z7github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pbb\x06proto3" var ( file_volume_server_proto_rawDescOnce sync.Once @@ -8431,7 +8617,7 @@ func file_volume_server_proto_rawDescGZIP() []byte { } var file_volume_server_proto_enumTypes = make([]protoimpl.EnumInfo, 1) -var file_volume_server_proto_msgTypes = make([]protoimpl.MessageInfo, 137) +var file_volume_server_proto_msgTypes = make([]protoimpl.MessageInfo, 139) var file_volume_server_proto_goTypes = []any{ (VolumeScrubMode)(0), // 0: volume_server_pb.VolumeScrubMode (*VolumeServerState)(nil), // 1: volume_server_pb.VolumeServerState @@ -8562,17 +8748,19 @@ var file_volume_server_proto_goTypes = []any{ (*CommitExpandBlockVolumeResponse)(nil), // 126: volume_server_pb.CommitExpandBlockVolumeResponse (*CancelExpandBlockVolumeRequest)(nil), // 127: volume_server_pb.CancelExpandBlockVolumeRequest (*CancelExpandBlockVolumeResponse)(nil), // 128: volume_server_pb.CancelExpandBlockVolumeResponse - (*FetchAndWriteNeedleRequest_Replica)(nil), // 129: volume_server_pb.FetchAndWriteNeedleRequest.Replica - (*QueryRequest_Filter)(nil), // 130: volume_server_pb.QueryRequest.Filter - (*QueryRequest_InputSerialization)(nil), // 131: volume_server_pb.QueryRequest.InputSerialization - (*QueryRequest_OutputSerialization)(nil), // 132: volume_server_pb.QueryRequest.OutputSerialization - (*QueryRequest_InputSerialization_CSVInput)(nil), // 133: volume_server_pb.QueryRequest.InputSerialization.CSVInput - (*QueryRequest_InputSerialization_JSONInput)(nil), // 134: volume_server_pb.QueryRequest.InputSerialization.JSONInput - (*QueryRequest_InputSerialization_ParquetInput)(nil), // 135: volume_server_pb.QueryRequest.InputSerialization.ParquetInput - (*QueryRequest_OutputSerialization_CSVOutput)(nil), // 136: volume_server_pb.QueryRequest.OutputSerialization.CSVOutput - (*QueryRequest_OutputSerialization_JSONOutput)(nil), // 137: volume_server_pb.QueryRequest.OutputSerialization.JSONOutput - (*remote_pb.RemoteConf)(nil), // 138: remote_pb.RemoteConf - (*remote_pb.RemoteStorageLocation)(nil), // 139: remote_pb.RemoteStorageLocation + (*QueryBlockPromotionEvidenceRequest)(nil), // 129: volume_server_pb.QueryBlockPromotionEvidenceRequest + (*QueryBlockPromotionEvidenceResponse)(nil), // 130: volume_server_pb.QueryBlockPromotionEvidenceResponse + (*FetchAndWriteNeedleRequest_Replica)(nil), // 131: volume_server_pb.FetchAndWriteNeedleRequest.Replica + (*QueryRequest_Filter)(nil), // 132: volume_server_pb.QueryRequest.Filter + (*QueryRequest_InputSerialization)(nil), // 133: volume_server_pb.QueryRequest.InputSerialization + (*QueryRequest_OutputSerialization)(nil), // 134: volume_server_pb.QueryRequest.OutputSerialization + (*QueryRequest_InputSerialization_CSVInput)(nil), // 135: volume_server_pb.QueryRequest.InputSerialization.CSVInput + (*QueryRequest_InputSerialization_JSONInput)(nil), // 136: volume_server_pb.QueryRequest.InputSerialization.JSONInput + (*QueryRequest_InputSerialization_ParquetInput)(nil), // 137: volume_server_pb.QueryRequest.InputSerialization.ParquetInput + (*QueryRequest_OutputSerialization_CSVOutput)(nil), // 138: volume_server_pb.QueryRequest.OutputSerialization.CSVOutput + (*QueryRequest_OutputSerialization_JSONOutput)(nil), // 139: volume_server_pb.QueryRequest.OutputSerialization.JSONOutput + (*remote_pb.RemoteConf)(nil), // 140: remote_pb.RemoteConf + (*remote_pb.RemoteStorageLocation)(nil), // 141: remote_pb.RemoteStorageLocation } var file_volume_server_proto_depIdxs = []int32{ 4, // 0: volume_server_pb.BatchDeleteResponse.results:type_name -> volume_server_pb.DeleteResult @@ -8588,21 +8776,21 @@ var file_volume_server_proto_depIdxs = []int32{ 82, // 10: volume_server_pb.VolumeServerStatusResponse.disk_statuses:type_name -> volume_server_pb.DiskStatus 83, // 11: volume_server_pb.VolumeServerStatusResponse.memory_status:type_name -> volume_server_pb.MemStatus 1, // 12: volume_server_pb.VolumeServerStatusResponse.state:type_name -> volume_server_pb.VolumeServerState - 129, // 13: volume_server_pb.FetchAndWriteNeedleRequest.replicas:type_name -> volume_server_pb.FetchAndWriteNeedleRequest.Replica - 138, // 14: volume_server_pb.FetchAndWriteNeedleRequest.remote_conf:type_name -> remote_pb.RemoteConf - 139, // 15: volume_server_pb.FetchAndWriteNeedleRequest.remote_location:type_name -> remote_pb.RemoteStorageLocation + 131, // 13: volume_server_pb.FetchAndWriteNeedleRequest.replicas:type_name -> volume_server_pb.FetchAndWriteNeedleRequest.Replica + 140, // 14: volume_server_pb.FetchAndWriteNeedleRequest.remote_conf:type_name -> remote_pb.RemoteConf + 141, // 15: volume_server_pb.FetchAndWriteNeedleRequest.remote_location:type_name -> remote_pb.RemoteStorageLocation 0, // 16: volume_server_pb.ScrubVolumeRequest.mode:type_name -> volume_server_pb.VolumeScrubMode 0, // 17: volume_server_pb.ScrubEcVolumeRequest.mode:type_name -> volume_server_pb.VolumeScrubMode 79, // 18: volume_server_pb.ScrubEcVolumeResponse.broken_shard_infos:type_name -> volume_server_pb.EcShardInfo - 130, // 19: volume_server_pb.QueryRequest.filter:type_name -> volume_server_pb.QueryRequest.Filter - 131, // 20: volume_server_pb.QueryRequest.input_serialization:type_name -> volume_server_pb.QueryRequest.InputSerialization - 132, // 21: volume_server_pb.QueryRequest.output_serialization:type_name -> volume_server_pb.QueryRequest.OutputSerialization + 132, // 19: volume_server_pb.QueryRequest.filter:type_name -> volume_server_pb.QueryRequest.Filter + 133, // 20: volume_server_pb.QueryRequest.input_serialization:type_name -> volume_server_pb.QueryRequest.InputSerialization + 134, // 21: volume_server_pb.QueryRequest.output_serialization:type_name -> volume_server_pb.QueryRequest.OutputSerialization 120, // 22: volume_server_pb.ListBlockSnapshotsResponse.snapshots:type_name -> volume_server_pb.BlockSnapshotInfo - 133, // 23: volume_server_pb.QueryRequest.InputSerialization.csv_input:type_name -> volume_server_pb.QueryRequest.InputSerialization.CSVInput - 134, // 24: volume_server_pb.QueryRequest.InputSerialization.json_input:type_name -> volume_server_pb.QueryRequest.InputSerialization.JSONInput - 135, // 25: volume_server_pb.QueryRequest.InputSerialization.parquet_input:type_name -> volume_server_pb.QueryRequest.InputSerialization.ParquetInput - 136, // 26: volume_server_pb.QueryRequest.OutputSerialization.csv_output:type_name -> volume_server_pb.QueryRequest.OutputSerialization.CSVOutput - 137, // 27: volume_server_pb.QueryRequest.OutputSerialization.json_output:type_name -> volume_server_pb.QueryRequest.OutputSerialization.JSONOutput + 135, // 23: volume_server_pb.QueryRequest.InputSerialization.csv_input:type_name -> volume_server_pb.QueryRequest.InputSerialization.CSVInput + 136, // 24: volume_server_pb.QueryRequest.InputSerialization.json_input:type_name -> volume_server_pb.QueryRequest.InputSerialization.JSONInput + 137, // 25: volume_server_pb.QueryRequest.InputSerialization.parquet_input:type_name -> volume_server_pb.QueryRequest.InputSerialization.ParquetInput + 138, // 26: volume_server_pb.QueryRequest.OutputSerialization.csv_output:type_name -> volume_server_pb.QueryRequest.OutputSerialization.CSVOutput + 139, // 27: volume_server_pb.QueryRequest.OutputSerialization.json_output:type_name -> volume_server_pb.QueryRequest.OutputSerialization.JSONOutput 2, // 28: volume_server_pb.VolumeServer.BatchDelete:input_type -> volume_server_pb.BatchDeleteRequest 6, // 29: volume_server_pb.VolumeServer.VacuumVolumeCheck:input_type -> volume_server_pb.VacuumVolumeCheckRequest 8, // 30: volume_server_pb.VolumeServer.VacuumVolumeCompact:input_type -> volume_server_pb.VacuumVolumeCompactRequest @@ -8661,66 +8849,68 @@ var file_volume_server_proto_depIdxs = []int32{ 123, // 83: volume_server_pb.VolumeServer.PrepareExpandBlockVolume:input_type -> volume_server_pb.PrepareExpandBlockVolumeRequest 125, // 84: volume_server_pb.VolumeServer.CommitExpandBlockVolume:input_type -> volume_server_pb.CommitExpandBlockVolumeRequest 127, // 85: volume_server_pb.VolumeServer.CancelExpandBlockVolume:input_type -> volume_server_pb.CancelExpandBlockVolumeRequest - 3, // 86: volume_server_pb.VolumeServer.BatchDelete:output_type -> volume_server_pb.BatchDeleteResponse - 7, // 87: volume_server_pb.VolumeServer.VacuumVolumeCheck:output_type -> volume_server_pb.VacuumVolumeCheckResponse - 9, // 88: volume_server_pb.VolumeServer.VacuumVolumeCompact:output_type -> volume_server_pb.VacuumVolumeCompactResponse - 11, // 89: volume_server_pb.VolumeServer.VacuumVolumeCommit:output_type -> volume_server_pb.VacuumVolumeCommitResponse - 13, // 90: volume_server_pb.VolumeServer.VacuumVolumeCleanup:output_type -> volume_server_pb.VacuumVolumeCleanupResponse - 15, // 91: volume_server_pb.VolumeServer.DeleteCollection:output_type -> volume_server_pb.DeleteCollectionResponse - 17, // 92: volume_server_pb.VolumeServer.AllocateVolume:output_type -> volume_server_pb.AllocateVolumeResponse - 19, // 93: volume_server_pb.VolumeServer.VolumeSyncStatus:output_type -> volume_server_pb.VolumeSyncStatusResponse - 21, // 94: volume_server_pb.VolumeServer.VolumeIncrementalCopy:output_type -> volume_server_pb.VolumeIncrementalCopyResponse - 23, // 95: volume_server_pb.VolumeServer.VolumeMount:output_type -> volume_server_pb.VolumeMountResponse - 25, // 96: volume_server_pb.VolumeServer.VolumeUnmount:output_type -> volume_server_pb.VolumeUnmountResponse - 27, // 97: volume_server_pb.VolumeServer.VolumeDelete:output_type -> volume_server_pb.VolumeDeleteResponse - 29, // 98: volume_server_pb.VolumeServer.VolumeMarkReadonly:output_type -> volume_server_pb.VolumeMarkReadonlyResponse - 31, // 99: volume_server_pb.VolumeServer.VolumeMarkWritable:output_type -> volume_server_pb.VolumeMarkWritableResponse - 33, // 100: volume_server_pb.VolumeServer.VolumeConfigure:output_type -> volume_server_pb.VolumeConfigureResponse - 35, // 101: volume_server_pb.VolumeServer.VolumeStatus:output_type -> volume_server_pb.VolumeStatusResponse - 37, // 102: volume_server_pb.VolumeServer.GetState:output_type -> volume_server_pb.GetStateResponse - 39, // 103: volume_server_pb.VolumeServer.SetState:output_type -> volume_server_pb.SetStateResponse - 41, // 104: volume_server_pb.VolumeServer.VolumeCopy:output_type -> volume_server_pb.VolumeCopyResponse - 81, // 105: volume_server_pb.VolumeServer.ReadVolumeFileStatus:output_type -> volume_server_pb.ReadVolumeFileStatusResponse - 43, // 106: volume_server_pb.VolumeServer.CopyFile:output_type -> volume_server_pb.CopyFileResponse - 46, // 107: volume_server_pb.VolumeServer.ReceiveFile:output_type -> volume_server_pb.ReceiveFileResponse - 48, // 108: volume_server_pb.VolumeServer.ReadNeedleBlob:output_type -> volume_server_pb.ReadNeedleBlobResponse - 50, // 109: volume_server_pb.VolumeServer.ReadNeedleMeta:output_type -> volume_server_pb.ReadNeedleMetaResponse - 52, // 110: volume_server_pb.VolumeServer.WriteNeedleBlob:output_type -> volume_server_pb.WriteNeedleBlobResponse - 54, // 111: volume_server_pb.VolumeServer.ReadAllNeedles:output_type -> volume_server_pb.ReadAllNeedlesResponse - 56, // 112: volume_server_pb.VolumeServer.VolumeTailSender:output_type -> volume_server_pb.VolumeTailSenderResponse - 58, // 113: volume_server_pb.VolumeServer.VolumeTailReceiver:output_type -> volume_server_pb.VolumeTailReceiverResponse - 60, // 114: volume_server_pb.VolumeServer.VolumeEcShardsGenerate:output_type -> volume_server_pb.VolumeEcShardsGenerateResponse - 62, // 115: volume_server_pb.VolumeServer.VolumeEcShardsRebuild:output_type -> volume_server_pb.VolumeEcShardsRebuildResponse - 64, // 116: volume_server_pb.VolumeServer.VolumeEcShardsCopy:output_type -> volume_server_pb.VolumeEcShardsCopyResponse - 66, // 117: volume_server_pb.VolumeServer.VolumeEcShardsDelete:output_type -> volume_server_pb.VolumeEcShardsDeleteResponse - 68, // 118: volume_server_pb.VolumeServer.VolumeEcShardsMount:output_type -> volume_server_pb.VolumeEcShardsMountResponse - 70, // 119: volume_server_pb.VolumeServer.VolumeEcShardsUnmount:output_type -> volume_server_pb.VolumeEcShardsUnmountResponse - 72, // 120: volume_server_pb.VolumeServer.VolumeEcShardRead:output_type -> volume_server_pb.VolumeEcShardReadResponse - 74, // 121: volume_server_pb.VolumeServer.VolumeEcBlobDelete:output_type -> volume_server_pb.VolumeEcBlobDeleteResponse - 76, // 122: volume_server_pb.VolumeServer.VolumeEcShardsToVolume:output_type -> volume_server_pb.VolumeEcShardsToVolumeResponse - 78, // 123: volume_server_pb.VolumeServer.VolumeEcShardsInfo:output_type -> volume_server_pb.VolumeEcShardsInfoResponse - 89, // 124: volume_server_pb.VolumeServer.VolumeTierMoveDatToRemote:output_type -> volume_server_pb.VolumeTierMoveDatToRemoteResponse - 91, // 125: volume_server_pb.VolumeServer.VolumeTierMoveDatFromRemote:output_type -> volume_server_pb.VolumeTierMoveDatFromRemoteResponse - 93, // 126: volume_server_pb.VolumeServer.VolumeServerStatus:output_type -> volume_server_pb.VolumeServerStatusResponse - 95, // 127: volume_server_pb.VolumeServer.VolumeServerLeave:output_type -> volume_server_pb.VolumeServerLeaveResponse - 97, // 128: volume_server_pb.VolumeServer.FetchAndWriteNeedle:output_type -> volume_server_pb.FetchAndWriteNeedleResponse - 99, // 129: volume_server_pb.VolumeServer.ScrubVolume:output_type -> volume_server_pb.ScrubVolumeResponse - 101, // 130: volume_server_pb.VolumeServer.ScrubEcVolume:output_type -> volume_server_pb.ScrubEcVolumeResponse - 103, // 131: volume_server_pb.VolumeServer.Query:output_type -> volume_server_pb.QueriedStripe - 105, // 132: volume_server_pb.VolumeServer.VolumeNeedleStatus:output_type -> volume_server_pb.VolumeNeedleStatusResponse - 107, // 133: volume_server_pb.VolumeServer.Ping:output_type -> volume_server_pb.PingResponse - 109, // 134: volume_server_pb.VolumeServer.AllocateBlockVolume:output_type -> volume_server_pb.AllocateBlockVolumeResponse - 111, // 135: volume_server_pb.VolumeServer.VolumeServerDeleteBlockVolume:output_type -> volume_server_pb.VolumeServerDeleteBlockVolumeResponse - 113, // 136: volume_server_pb.VolumeServer.SnapshotBlockVolume:output_type -> volume_server_pb.SnapshotBlockVolumeResponse - 115, // 137: volume_server_pb.VolumeServer.DeleteBlockSnapshot:output_type -> volume_server_pb.DeleteBlockSnapshotResponse - 119, // 138: volume_server_pb.VolumeServer.ListBlockSnapshots:output_type -> volume_server_pb.ListBlockSnapshotsResponse - 117, // 139: volume_server_pb.VolumeServer.RestoreBlockSnapshot:output_type -> volume_server_pb.RestoreBlockSnapshotResponse - 122, // 140: volume_server_pb.VolumeServer.ExpandBlockVolume:output_type -> volume_server_pb.ExpandBlockVolumeResponse - 124, // 141: volume_server_pb.VolumeServer.PrepareExpandBlockVolume:output_type -> volume_server_pb.PrepareExpandBlockVolumeResponse - 126, // 142: volume_server_pb.VolumeServer.CommitExpandBlockVolume:output_type -> volume_server_pb.CommitExpandBlockVolumeResponse - 128, // 143: volume_server_pb.VolumeServer.CancelExpandBlockVolume:output_type -> volume_server_pb.CancelExpandBlockVolumeResponse - 86, // [86:144] is the sub-list for method output_type - 28, // [28:86] is the sub-list for method input_type + 129, // 86: volume_server_pb.VolumeServer.QueryBlockPromotionEvidence:input_type -> volume_server_pb.QueryBlockPromotionEvidenceRequest + 3, // 87: volume_server_pb.VolumeServer.BatchDelete:output_type -> volume_server_pb.BatchDeleteResponse + 7, // 88: volume_server_pb.VolumeServer.VacuumVolumeCheck:output_type -> volume_server_pb.VacuumVolumeCheckResponse + 9, // 89: volume_server_pb.VolumeServer.VacuumVolumeCompact:output_type -> volume_server_pb.VacuumVolumeCompactResponse + 11, // 90: volume_server_pb.VolumeServer.VacuumVolumeCommit:output_type -> volume_server_pb.VacuumVolumeCommitResponse + 13, // 91: volume_server_pb.VolumeServer.VacuumVolumeCleanup:output_type -> volume_server_pb.VacuumVolumeCleanupResponse + 15, // 92: volume_server_pb.VolumeServer.DeleteCollection:output_type -> volume_server_pb.DeleteCollectionResponse + 17, // 93: volume_server_pb.VolumeServer.AllocateVolume:output_type -> volume_server_pb.AllocateVolumeResponse + 19, // 94: volume_server_pb.VolumeServer.VolumeSyncStatus:output_type -> volume_server_pb.VolumeSyncStatusResponse + 21, // 95: volume_server_pb.VolumeServer.VolumeIncrementalCopy:output_type -> volume_server_pb.VolumeIncrementalCopyResponse + 23, // 96: volume_server_pb.VolumeServer.VolumeMount:output_type -> volume_server_pb.VolumeMountResponse + 25, // 97: volume_server_pb.VolumeServer.VolumeUnmount:output_type -> volume_server_pb.VolumeUnmountResponse + 27, // 98: volume_server_pb.VolumeServer.VolumeDelete:output_type -> volume_server_pb.VolumeDeleteResponse + 29, // 99: volume_server_pb.VolumeServer.VolumeMarkReadonly:output_type -> volume_server_pb.VolumeMarkReadonlyResponse + 31, // 100: volume_server_pb.VolumeServer.VolumeMarkWritable:output_type -> volume_server_pb.VolumeMarkWritableResponse + 33, // 101: volume_server_pb.VolumeServer.VolumeConfigure:output_type -> volume_server_pb.VolumeConfigureResponse + 35, // 102: volume_server_pb.VolumeServer.VolumeStatus:output_type -> volume_server_pb.VolumeStatusResponse + 37, // 103: volume_server_pb.VolumeServer.GetState:output_type -> volume_server_pb.GetStateResponse + 39, // 104: volume_server_pb.VolumeServer.SetState:output_type -> volume_server_pb.SetStateResponse + 41, // 105: volume_server_pb.VolumeServer.VolumeCopy:output_type -> volume_server_pb.VolumeCopyResponse + 81, // 106: volume_server_pb.VolumeServer.ReadVolumeFileStatus:output_type -> volume_server_pb.ReadVolumeFileStatusResponse + 43, // 107: volume_server_pb.VolumeServer.CopyFile:output_type -> volume_server_pb.CopyFileResponse + 46, // 108: volume_server_pb.VolumeServer.ReceiveFile:output_type -> volume_server_pb.ReceiveFileResponse + 48, // 109: volume_server_pb.VolumeServer.ReadNeedleBlob:output_type -> volume_server_pb.ReadNeedleBlobResponse + 50, // 110: volume_server_pb.VolumeServer.ReadNeedleMeta:output_type -> volume_server_pb.ReadNeedleMetaResponse + 52, // 111: volume_server_pb.VolumeServer.WriteNeedleBlob:output_type -> volume_server_pb.WriteNeedleBlobResponse + 54, // 112: volume_server_pb.VolumeServer.ReadAllNeedles:output_type -> volume_server_pb.ReadAllNeedlesResponse + 56, // 113: volume_server_pb.VolumeServer.VolumeTailSender:output_type -> volume_server_pb.VolumeTailSenderResponse + 58, // 114: volume_server_pb.VolumeServer.VolumeTailReceiver:output_type -> volume_server_pb.VolumeTailReceiverResponse + 60, // 115: volume_server_pb.VolumeServer.VolumeEcShardsGenerate:output_type -> volume_server_pb.VolumeEcShardsGenerateResponse + 62, // 116: volume_server_pb.VolumeServer.VolumeEcShardsRebuild:output_type -> volume_server_pb.VolumeEcShardsRebuildResponse + 64, // 117: volume_server_pb.VolumeServer.VolumeEcShardsCopy:output_type -> volume_server_pb.VolumeEcShardsCopyResponse + 66, // 118: volume_server_pb.VolumeServer.VolumeEcShardsDelete:output_type -> volume_server_pb.VolumeEcShardsDeleteResponse + 68, // 119: volume_server_pb.VolumeServer.VolumeEcShardsMount:output_type -> volume_server_pb.VolumeEcShardsMountResponse + 70, // 120: volume_server_pb.VolumeServer.VolumeEcShardsUnmount:output_type -> volume_server_pb.VolumeEcShardsUnmountResponse + 72, // 121: volume_server_pb.VolumeServer.VolumeEcShardRead:output_type -> volume_server_pb.VolumeEcShardReadResponse + 74, // 122: volume_server_pb.VolumeServer.VolumeEcBlobDelete:output_type -> volume_server_pb.VolumeEcBlobDeleteResponse + 76, // 123: volume_server_pb.VolumeServer.VolumeEcShardsToVolume:output_type -> volume_server_pb.VolumeEcShardsToVolumeResponse + 78, // 124: volume_server_pb.VolumeServer.VolumeEcShardsInfo:output_type -> volume_server_pb.VolumeEcShardsInfoResponse + 89, // 125: volume_server_pb.VolumeServer.VolumeTierMoveDatToRemote:output_type -> volume_server_pb.VolumeTierMoveDatToRemoteResponse + 91, // 126: volume_server_pb.VolumeServer.VolumeTierMoveDatFromRemote:output_type -> volume_server_pb.VolumeTierMoveDatFromRemoteResponse + 93, // 127: volume_server_pb.VolumeServer.VolumeServerStatus:output_type -> volume_server_pb.VolumeServerStatusResponse + 95, // 128: volume_server_pb.VolumeServer.VolumeServerLeave:output_type -> volume_server_pb.VolumeServerLeaveResponse + 97, // 129: volume_server_pb.VolumeServer.FetchAndWriteNeedle:output_type -> volume_server_pb.FetchAndWriteNeedleResponse + 99, // 130: volume_server_pb.VolumeServer.ScrubVolume:output_type -> volume_server_pb.ScrubVolumeResponse + 101, // 131: volume_server_pb.VolumeServer.ScrubEcVolume:output_type -> volume_server_pb.ScrubEcVolumeResponse + 103, // 132: volume_server_pb.VolumeServer.Query:output_type -> volume_server_pb.QueriedStripe + 105, // 133: volume_server_pb.VolumeServer.VolumeNeedleStatus:output_type -> volume_server_pb.VolumeNeedleStatusResponse + 107, // 134: volume_server_pb.VolumeServer.Ping:output_type -> volume_server_pb.PingResponse + 109, // 135: volume_server_pb.VolumeServer.AllocateBlockVolume:output_type -> volume_server_pb.AllocateBlockVolumeResponse + 111, // 136: volume_server_pb.VolumeServer.VolumeServerDeleteBlockVolume:output_type -> volume_server_pb.VolumeServerDeleteBlockVolumeResponse + 113, // 137: volume_server_pb.VolumeServer.SnapshotBlockVolume:output_type -> volume_server_pb.SnapshotBlockVolumeResponse + 115, // 138: volume_server_pb.VolumeServer.DeleteBlockSnapshot:output_type -> volume_server_pb.DeleteBlockSnapshotResponse + 119, // 139: volume_server_pb.VolumeServer.ListBlockSnapshots:output_type -> volume_server_pb.ListBlockSnapshotsResponse + 117, // 140: volume_server_pb.VolumeServer.RestoreBlockSnapshot:output_type -> volume_server_pb.RestoreBlockSnapshotResponse + 122, // 141: volume_server_pb.VolumeServer.ExpandBlockVolume:output_type -> volume_server_pb.ExpandBlockVolumeResponse + 124, // 142: volume_server_pb.VolumeServer.PrepareExpandBlockVolume:output_type -> volume_server_pb.PrepareExpandBlockVolumeResponse + 126, // 143: volume_server_pb.VolumeServer.CommitExpandBlockVolume:output_type -> volume_server_pb.CommitExpandBlockVolumeResponse + 128, // 144: volume_server_pb.VolumeServer.CancelExpandBlockVolume:output_type -> volume_server_pb.CancelExpandBlockVolumeResponse + 130, // 145: volume_server_pb.VolumeServer.QueryBlockPromotionEvidence:output_type -> volume_server_pb.QueryBlockPromotionEvidenceResponse + 87, // [87:146] is the sub-list for method output_type + 28, // [28:87] is the sub-list for method input_type 28, // [28:28] is the sub-list for extension type_name 28, // [28:28] is the sub-list for extension extendee 0, // [0:28] is the sub-list for field type_name @@ -8741,7 +8931,7 @@ func file_volume_server_proto_init() { GoPackagePath: reflect.TypeOf(x{}).PkgPath(), RawDescriptor: unsafe.Slice(unsafe.StringData(file_volume_server_proto_rawDesc), len(file_volume_server_proto_rawDesc)), NumEnums: 1, - NumMessages: 137, + NumMessages: 139, NumExtensions: 0, NumServices: 1, }, diff --git a/weed/pb/volume_server_pb/volume_server_grpc.pb.go b/weed/pb/volume_server_pb/volume_server_grpc.pb.go index 85aa6d176..5701ddb6f 100644 --- a/weed/pb/volume_server_pb/volume_server_grpc.pb.go +++ b/weed/pb/volume_server_pb/volume_server_grpc.pb.go @@ -77,6 +77,7 @@ const ( VolumeServer_PrepareExpandBlockVolume_FullMethodName = "/volume_server_pb.VolumeServer/PrepareExpandBlockVolume" VolumeServer_CommitExpandBlockVolume_FullMethodName = "/volume_server_pb.VolumeServer/CommitExpandBlockVolume" VolumeServer_CancelExpandBlockVolume_FullMethodName = "/volume_server_pb.VolumeServer/CancelExpandBlockVolume" + VolumeServer_QueryBlockPromotionEvidence_FullMethodName = "/volume_server_pb.VolumeServer/QueryBlockPromotionEvidence" ) // VolumeServerClient is the client API for VolumeServer service. @@ -149,6 +150,7 @@ type VolumeServerClient interface { PrepareExpandBlockVolume(ctx context.Context, in *PrepareExpandBlockVolumeRequest, opts ...grpc.CallOption) (*PrepareExpandBlockVolumeResponse, error) CommitExpandBlockVolume(ctx context.Context, in *CommitExpandBlockVolumeRequest, opts ...grpc.CallOption) (*CommitExpandBlockVolumeResponse, error) CancelExpandBlockVolume(ctx context.Context, in *CancelExpandBlockVolumeRequest, opts ...grpc.CallOption) (*CancelExpandBlockVolumeResponse, error) + QueryBlockPromotionEvidence(ctx context.Context, in *QueryBlockPromotionEvidenceRequest, opts ...grpc.CallOption) (*QueryBlockPromotionEvidenceResponse, error) } type volumeServerClient struct { @@ -832,6 +834,16 @@ func (c *volumeServerClient) CancelExpandBlockVolume(ctx context.Context, in *Ca return out, nil } +func (c *volumeServerClient) QueryBlockPromotionEvidence(ctx context.Context, in *QueryBlockPromotionEvidenceRequest, opts ...grpc.CallOption) (*QueryBlockPromotionEvidenceResponse, error) { + cOpts := append([]grpc.CallOption{grpc.StaticMethod()}, opts...) + out := new(QueryBlockPromotionEvidenceResponse) + err := c.cc.Invoke(ctx, VolumeServer_QueryBlockPromotionEvidence_FullMethodName, in, out, cOpts...) + if err != nil { + return nil, err + } + return out, nil +} + // VolumeServerServer is the server API for VolumeServer service. // All implementations must embed UnimplementedVolumeServerServer // for forward compatibility. @@ -902,6 +914,7 @@ type VolumeServerServer interface { PrepareExpandBlockVolume(context.Context, *PrepareExpandBlockVolumeRequest) (*PrepareExpandBlockVolumeResponse, error) CommitExpandBlockVolume(context.Context, *CommitExpandBlockVolumeRequest) (*CommitExpandBlockVolumeResponse, error) CancelExpandBlockVolume(context.Context, *CancelExpandBlockVolumeRequest) (*CancelExpandBlockVolumeResponse, error) + QueryBlockPromotionEvidence(context.Context, *QueryBlockPromotionEvidenceRequest) (*QueryBlockPromotionEvidenceResponse, error) mustEmbedUnimplementedVolumeServerServer() } @@ -1086,6 +1099,9 @@ func (UnimplementedVolumeServerServer) CommitExpandBlockVolume(context.Context, func (UnimplementedVolumeServerServer) CancelExpandBlockVolume(context.Context, *CancelExpandBlockVolumeRequest) (*CancelExpandBlockVolumeResponse, error) { return nil, status.Error(codes.Unimplemented, "method CancelExpandBlockVolume not implemented") } +func (UnimplementedVolumeServerServer) QueryBlockPromotionEvidence(context.Context, *QueryBlockPromotionEvidenceRequest) (*QueryBlockPromotionEvidenceResponse, error) { + return nil, status.Error(codes.Unimplemented, "method QueryBlockPromotionEvidence not implemented") +} func (UnimplementedVolumeServerServer) mustEmbedUnimplementedVolumeServerServer() {} func (UnimplementedVolumeServerServer) testEmbeddedByValue() {} @@ -2070,6 +2086,24 @@ func _VolumeServer_CancelExpandBlockVolume_Handler(srv interface{}, ctx context. return interceptor(ctx, in, info, handler) } +func _VolumeServer_QueryBlockPromotionEvidence_Handler(srv interface{}, ctx context.Context, dec func(interface{}) error, interceptor grpc.UnaryServerInterceptor) (interface{}, error) { + in := new(QueryBlockPromotionEvidenceRequest) + if err := dec(in); err != nil { + return nil, err + } + if interceptor == nil { + return srv.(VolumeServerServer).QueryBlockPromotionEvidence(ctx, in) + } + info := &grpc.UnaryServerInfo{ + Server: srv, + FullMethod: VolumeServer_QueryBlockPromotionEvidence_FullMethodName, + } + handler := func(ctx context.Context, req interface{}) (interface{}, error) { + return srv.(VolumeServerServer).QueryBlockPromotionEvidence(ctx, req.(*QueryBlockPromotionEvidenceRequest)) + } + return interceptor(ctx, in, info, handler) +} + // VolumeServer_ServiceDesc is the grpc.ServiceDesc for VolumeServer service. // It's only intended for direct use with grpc.RegisterService, // and not to be introspected or modified (even as a copy) @@ -2265,6 +2299,10 @@ var VolumeServer_ServiceDesc = grpc.ServiceDesc{ MethodName: "CancelExpandBlockVolume", Handler: _VolumeServer_CancelExpandBlockVolume_Handler, }, + { + MethodName: "QueryBlockPromotionEvidence", + Handler: _VolumeServer_QueryBlockPromotionEvidence_Handler, + }, }, Streams: []grpc.StreamDesc{ { diff --git a/weed/server/integration_block_test.go b/weed/server/integration_block_test.go index 2517cd68d..c437fddf9 100644 --- a/weed/server/integration_block_test.go +++ b/weed/server/integration_block_test.go @@ -28,7 +28,7 @@ func integrationMaster(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -385,13 +385,13 @@ func TestIntegration_ReplicaFailureSingleCopy(t *testing.T) { // Make replica allocation always fail. callCount := 0 origAllocate := ms.blockVSAllocate - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount++ if callCount > 1 { // Second call (replica) fails. return nil, fmt.Errorf("disk full on replica") } - return origAllocate(ctx, server, name, sizeBytes, diskType, durabilityMode) + return origAllocate(ctx, server, name, sizeBytes, walSizeBytes, diskType, durabilityMode) } resp, err := ms.CreateBlockVolume(ctx, &master_pb.CreateBlockVolumeRequest{ diff --git a/weed/server/master_block_failover_test.go b/weed/server/master_block_failover_test.go index 037f12836..4aa452160 100644 --- a/weed/server/master_block_failover_test.go +++ b/weed/server/master_block_failover_test.go @@ -18,7 +18,7 @@ func testMasterServerForFailover(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), diff --git a/weed/server/master_block_plan_test.go b/weed/server/master_block_plan_test.go index 3183f3875..64402e9d2 100644 --- a/weed/server/master_block_plan_test.go +++ b/weed/server/master_block_plan_test.go @@ -266,7 +266,7 @@ func qaPlanMaster(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), diff --git a/weed/server/master_grpc_server_block.go b/weed/server/master_grpc_server_block.go index d650fedd9..3be6bf370 100644 --- a/weed/server/master_grpc_server_block.go +++ b/weed/server/master_grpc_server_block.go @@ -86,7 +86,7 @@ func (ms *MasterServer) CreateBlockVolume(ctx context.Context, req *master_pb.Cr for attempt := 0; attempt < len(placement.Candidates); attempt++ { server := placement.Candidates[attempt] - result, err := ms.blockVSAllocate(ctx, server, req.Name, req.SizeBytes, req.DiskType, req.DurabilityMode) + result, err := ms.blockVSAllocate(ctx, server, req.Name, req.SizeBytes, req.WalSizeBytes, req.DiskType, req.DurabilityMode) if err != nil { lastErr = fmt.Errorf("server %s: %w", server, err) glog.V(0).Infof("[reqID=%s] CreateBlockVolume %q: attempt %d on %s failed: %v", blockReqID(ctx), req.Name, attempt+1, server, err) @@ -281,7 +281,7 @@ func lookupResponseFromEntry(entry *BlockVolumeEntry) *master_pb.LookupBlockVolu // Returns the replica server address on success, or empty string on failure (F4). func (ms *MasterServer) tryCreateOneReplica(ctx context.Context, req *master_pb.CreateBlockVolumeRequest, entry *BlockVolumeEntry, primaryResult *blockAllocResult, candidates []string) string { for _, replicaServerStr := range candidates { - replicaResult, err := ms.blockVSAllocate(ctx, replicaServerStr, req.Name, req.SizeBytes, req.DiskType, req.DurabilityMode) + replicaResult, err := ms.blockVSAllocate(ctx, replicaServerStr, req.Name, req.SizeBytes, req.WalSizeBytes, req.DiskType, req.DurabilityMode) if err != nil { glog.V(0).Infof("[reqID=%s] CreateBlockVolume %q: replica on %s failed: %v", blockReqID(ctx), req.Name, replicaServerStr, err) continue diff --git a/weed/server/master_grpc_server_block_test.go b/weed/server/master_grpc_server_block_test.go index 3893dd97d..52f88c5e8 100644 --- a/weed/server/master_grpc_server_block_test.go +++ b/weed/server/master_grpc_server_block_test.go @@ -21,7 +21,7 @@ func testMasterServer(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), } // Default mock: succeed with deterministic values. - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -141,7 +141,7 @@ func TestMaster_CreateVSFailure_Retry(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs2:9333") var callCount atomic.Int32 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { n := callCount.Add(1) if n == 1 { return nil, fmt.Errorf("disk full") @@ -172,7 +172,7 @@ func TestMaster_CreateVSFailure_Cleanup(t *testing.T) { ms := testMasterServer(t) ms.blockRegistry.MarkBlockCapable("vs1:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return nil, fmt.Errorf("all servers broken") } @@ -195,7 +195,7 @@ func TestMaster_CreateConcurrentSameName(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs1:9333") var callCount atomic.Int32 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount.Add(1) return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), @@ -277,7 +277,7 @@ func TestMaster_CreateWithReplica(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs2:9333") var allocServers []string - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { allocServers = append(allocServers, server) return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), @@ -330,7 +330,7 @@ func TestMaster_CreateSingleServer_NoReplica(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs1:9333") var allocCount atomic.Int32 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { allocCount.Add(1) return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), @@ -367,7 +367,7 @@ func TestMaster_CreateReplica_SecondFails_SingleCopy(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs2:9333") var callCount atomic.Int32 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { n := callCount.Add(1) if n == 2 { // Replica allocation fails. @@ -404,7 +404,7 @@ func TestMaster_CreateEnqueuesAssignments(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -465,7 +465,7 @@ func TestMaster_LookupReturnsReplicaServer(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -713,7 +713,7 @@ func TestLookupResponseFromEntry_PublicationMinimalSurface(t *testing.T) { func testMasterServerRF3(t *testing.T) *MasterServer { t.Helper() ms := testMasterServer(t) - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -768,7 +768,7 @@ func TestMaster_CreateRF3_ThreeServers(t *testing.T) { // RF=3 with only 2 servers: should create 1 replica (partial). func TestMaster_CreateRF3_TwoServers(t *testing.T) { ms := testMasterServer(t) - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -1071,7 +1071,7 @@ func TestMaster_NvmeFieldsFlowThroughCreateAndLookup(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs1:9333") // Mock: VS returns NVMe fields. - ms.blockVSAllocate = func(ctx context.Context, server, name string, sizeBytes uint64, diskType, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server, name string, sizeBytes uint64, walSizeBytes uint64, diskType, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -1261,7 +1261,7 @@ func TestMaster_ExpandCoordinated_Success(t *testing.T) { ms := testMasterServerWithExpandMocks(t) ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.test:%s", name), @@ -1310,7 +1310,7 @@ func TestMaster_ExpandCoordinated_PrepareFailure_Cancels(t *testing.T) { ms := testMasterServerWithExpandMocks(t) ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.test:%s", name), @@ -1387,7 +1387,7 @@ func TestMaster_ExpandCoordinated_ConcurrentRejected(t *testing.T) { ms := testMasterServerWithExpandMocks(t) ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.test:%s", name), @@ -1461,7 +1461,7 @@ func TestMaster_ExpandCoordinated_CommitFailure_MarksInconsistent(t *testing.T) ms := testMasterServerWithExpandMocks(t) ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.test:%s", name), @@ -1542,7 +1542,7 @@ func TestMaster_ExpandCoordinated_HeartbeatSuppressedAfterPartialCommit(t *testi ms := testMasterServerWithExpandMocks(t) ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.test:%s", name), @@ -1602,7 +1602,7 @@ func TestMaster_ExpandCoordinated_FailoverDuringPrepare(t *testing.T) { ms := testMasterServerWithExpandMocks(t) ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.test:%s", name), @@ -1655,7 +1655,7 @@ func TestMaster_ExpandCoordinated_RestartRecovery(t *testing.T) { ms := testMasterServerWithExpandMocks(t) ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.test:%s", name), @@ -1710,7 +1710,7 @@ func TestMaster_ExpandCoordinated_B09_ReReadsEntryAfterLock(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") ms.blockRegistry.MarkBlockCapable("vs3:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.test:%s", name), @@ -1791,7 +1791,7 @@ func TestMaster_ExpandCoordinated_B10_HeartbeatDoesNotDeleteDuringExpand(t *test ms := testMasterServerWithExpandMocks(t) ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.test:%s", name), diff --git a/weed/server/master_server.go b/weed/server/master_server.go index 0fbb775c4..6ea4668ac 100644 --- a/weed/server/master_server.go +++ b/weed/server/master_server.go @@ -54,18 +54,18 @@ type MasterOption struct { VolumePreallocate bool MaxParallelVacuumPerServer int // PulseSeconds int - DefaultReplicaPlacement string - GarbageThreshold float64 - WhiteList []string - DisableHttp bool - MetricsAddress string - MetricsIntervalSec int - IsFollower bool - TelemetryUrl string - TelemetryEnabled bool - VolumeGrowthDisabled bool - BlockPromotionLSNTolerance int - BlockV2Promotion bool // T3: enable durability-first V2 promotion + DefaultReplicaPlacement string + GarbageThreshold float64 + WhiteList []string + DisableHttp bool + MetricsAddress string + MetricsIntervalSec int + IsFollower bool + TelemetryUrl string + TelemetryEnabled bool + VolumeGrowthDisabled bool + BlockPromotionLSNTolerance int + BlockV2Promotion bool // T3: enable durability-first V2 promotion } type MasterServer struct { @@ -97,23 +97,23 @@ type MasterServer struct { telemetryCollector *telemetry.Collector // block volume support - blockRegistry *BlockVolumeRegistry - blockAssignmentQueue *BlockAssignmentQueue - blockFailover *blockFailoverState - blockVSAllocate func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) - blockVSDelete func(ctx context.Context, server string, name string) error - blockVSSnapshot func(ctx context.Context, server string, name string, snapID uint32) (int64, uint64, error) - blockVSDeleteSnap func(ctx context.Context, server string, name string, snapID uint32) error - blockVSListSnaps func(ctx context.Context, server string, name string) ([]*volume_server_pb.BlockSnapshotInfo, error) - blockVSRestore func(ctx context.Context, server string, name string, snapID uint32) error - blockVSExpand func(ctx context.Context, server string, name string, newSize uint64) (uint64, error) - blockVSPrepareExpand func(ctx context.Context, server string, name string, newSize, expandEpoch uint64) error - blockVSCommitExpand func(ctx context.Context, server string, name string, expandEpoch uint64) (uint64, error) - blockVSCancelExpand func(ctx context.Context, server string, name string, expandEpoch uint64) error - blockVSQueryEvidence BlockPromotionEvidenceQuerier // T2: fresh on-demand promotion evidence - blockV2EvidenceTransport bool // T3: true only when real gRPC querier is installed (not placeholder) - nextExpandEpoch atomic.Uint64 - blockV2Promotion bool // T3: when true, use durability-first V2 promotion; when false, legacy V1 + blockRegistry *BlockVolumeRegistry + blockAssignmentQueue *BlockAssignmentQueue + blockFailover *blockFailoverState + blockVSAllocate func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) + blockVSDelete func(ctx context.Context, server string, name string) error + blockVSSnapshot func(ctx context.Context, server string, name string, snapID uint32) (int64, uint64, error) + blockVSDeleteSnap func(ctx context.Context, server string, name string, snapID uint32) error + blockVSListSnaps func(ctx context.Context, server string, name string) ([]*volume_server_pb.BlockSnapshotInfo, error) + blockVSRestore func(ctx context.Context, server string, name string, snapID uint32) error + blockVSExpand func(ctx context.Context, server string, name string, newSize uint64) (uint64, error) + blockVSPrepareExpand func(ctx context.Context, server string, name string, newSize, expandEpoch uint64) error + blockVSCommitExpand func(ctx context.Context, server string, name string, expandEpoch uint64) (uint64, error) + blockVSCancelExpand func(ctx context.Context, server string, name string, expandEpoch uint64) error + blockVSQueryEvidence BlockPromotionEvidenceQuerier // T2: fresh on-demand promotion evidence + blockV2EvidenceTransport bool // T3: true only when real gRPC querier is installed (not placeholder) + nextExpandEpoch atomic.Uint64 + blockV2Promotion bool // T3: when true, use durability-first V2 promotion; when false, legacy V1 // Test-only hook: called after AcquireExpandInflight but before the // re-read Lookup in coordinated expand. Nil in production. @@ -571,23 +571,24 @@ func (ms *MasterServer) Reload() { // blockAllocResult holds the result of a block volume allocation. type blockAllocResult struct { - Path string - IQN string - ISCSIAddr string - ReplicaDataAddr string - ReplicaCtrlAddr string + Path string + IQN string + ISCSIAddr string + ReplicaDataAddr string + ReplicaCtrlAddr string RebuildListenAddr string - NvmeAddr string - NQN string + NvmeAddr string + NQN string } // defaultBlockVSAllocate calls a volume server's AllocateBlockVolume RPC. -func (ms *MasterServer) defaultBlockVSAllocate(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { +func (ms *MasterServer) defaultBlockVSAllocate(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { var result blockAllocResult err := operation.WithVolumeServerClient(false, pb.ServerAddress(server), ms.grpcDialOption, func(client volume_server_pb.VolumeServerClient) error { resp, rerr := client.AllocateBlockVolume(ctx, &volume_server_pb.AllocateBlockVolumeRequest{ Name: name, SizeBytes: sizeBytes, + WalSizeBytes: walSizeBytes, DiskType: diskType, DurabilityMode: durabilityMode, }) diff --git a/weed/server/master_server_handlers_block.go b/weed/server/master_server_handlers_block.go index 1ea32c76c..be87a8fdc 100644 --- a/weed/server/master_server_handlers_block.go +++ b/weed/server/master_server_handlers_block.go @@ -50,6 +50,7 @@ func (ms *MasterServer) blockVolumeCreateHandler(w http.ResponseWriter, r *http. resp, err := ms.CreateBlockVolume(r.Context(), &master_pb.CreateBlockVolumeRequest{ Name: req.Name, SizeBytes: req.SizeBytes, + WalSizeBytes: req.WALSizeBytes, DiskType: resolved.Policy.DiskType, DurabilityMode: resolved.Policy.DurabilityMode, ReplicaFactor: uint32(resolved.Policy.ReplicaFactor), @@ -409,29 +410,29 @@ func entryToVolumeInfo(e *BlockVolumeEntry, primaryAlive bool) blockapi.VolumeIn } surface := entryReplicaSurfaceInfo(e, primaryAlive) info := blockapi.VolumeInfo{ - Name: e.Name, - VolumeServer: e.VolumeServer, - SizeBytes: e.SizeBytes, - ReplicaPlacement: e.ReplicaPlacement, - Epoch: e.Epoch, - Role: blockvol.RoleFromWire(e.Role).String(), - Status: status, - ISCSIAddr: e.ISCSIAddr, - IQN: e.IQN, - ReplicaServer: e.ReplicaServer, - ReplicaISCSIAddr: e.ReplicaISCSIAddr, - ReplicaIQN: e.ReplicaIQN, - ReplicaDataAddr: e.ReplicaDataAddr, - ReplicaCtrlAddr: e.ReplicaCtrlAddr, - ReplicaFactor: rf, - ReplicaReady: surface.ReplicaReady, - HealthScore: e.HealthScore, - ReplicaDegraded: surface.ReplicaDegraded, - DurabilityMode: durMode, - Preset: e.Preset, - NvmeAddr: e.NvmeAddr, - NQN: e.NQN, - HealthState: surface.HealthState, + Name: e.Name, + VolumeServer: e.VolumeServer, + SizeBytes: e.SizeBytes, + ReplicaPlacement: e.ReplicaPlacement, + Epoch: e.Epoch, + Role: blockvol.RoleFromWire(e.Role).String(), + Status: status, + ISCSIAddr: e.ISCSIAddr, + IQN: e.IQN, + ReplicaServer: e.ReplicaServer, + ReplicaISCSIAddr: e.ReplicaISCSIAddr, + ReplicaIQN: e.ReplicaIQN, + ReplicaDataAddr: e.ReplicaDataAddr, + ReplicaCtrlAddr: e.ReplicaCtrlAddr, + ReplicaFactor: rf, + ReplicaReady: surface.ReplicaReady, + HealthScore: e.HealthScore, + ReplicaDegraded: surface.ReplicaDegraded, + DurabilityMode: durMode, + Preset: e.Preset, + NvmeAddr: e.NvmeAddr, + NQN: e.NQN, + HealthState: surface.HealthState, VolumeMode: surface.VolumeMode, VolumeModeReason: surface.VolumeModeReason, EngineProjectionMode: e.EngineProjectionMode, diff --git a/weed/server/master_server_handlers_block_test.go b/weed/server/master_server_handlers_block_test.go index c7b2fa263..2e308bf00 100644 --- a/weed/server/master_server_handlers_block_test.go +++ b/weed/server/master_server_handlers_block_test.go @@ -22,7 +22,7 @@ func blockTestServer(t *testing.T) (*MasterServer, *httptest.Server) { blockRegistry: NewBlockVolumeRegistry(), blockAssignmentQueue: NewBlockAssignmentQueue(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -75,6 +75,37 @@ func TestBlockVolumeCreateHandler(t *testing.T) { } } +func TestBlockVolumeCreateHandler_ForwardsWALSizeBytes(t *testing.T) { + ms, ts := blockTestServer(t) + + var capturedWalSize uint64 + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + capturedWalSize = walSizeBytes + return &blockAllocResult{ + Path: fmt.Sprintf("/data/%s.blk", name), + IQN: fmt.Sprintf("iqn.2024.test:%s", name), + ISCSIAddr: server + ":3260", + }, nil + } + + body, _ := json.Marshal(blockapi.CreateVolumeRequest{ + Name: "vol-wal", + SizeBytes: 1 << 30, + WALSizeBytes: 256 << 20, + }) + resp, err := http.Post(ts.URL+"/block/volume", "application/json", bytes.NewReader(body)) + if err != nil { + t.Fatal(err) + } + defer resp.Body.Close() + if resp.StatusCode != http.StatusOK { + t.Fatalf("expected 200, got %d", resp.StatusCode) + } + if capturedWalSize != 256<<20 { + t.Fatalf("wal_size_bytes=%d, want %d", capturedWalSize, 256<<20) + } +} + func TestBlockVolumeListHandler(t *testing.T) { ms, ts := blockTestServer(t) diff --git a/weed/server/qa_block_control_loop_test.go b/weed/server/qa_block_control_loop_test.go index 2bb472e79..c7148bf26 100644 --- a/weed/server/qa_block_control_loop_test.go +++ b/weed/server/qa_block_control_loop_test.go @@ -61,7 +61,7 @@ func newP4Setup(t *testing.T) *p4Setup { setup := &p4Setup{ms: ms, bs: bs, store: store, dir: dir} // Wire allocator to create REAL block volumes at per-server paths. - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { // Per-server subdir (sanitize colons for Windows). sanitized := strings.ReplaceAll(server, ":", "_") serverDir := filepath.Join(dir, sanitized) diff --git a/weed/server/qa_block_cp11b1_adversarial_test.go b/weed/server/qa_block_cp11b1_adversarial_test.go index 94c480143..1d77c8ba0 100644 --- a/weed/server/qa_block_cp11b1_adversarial_test.go +++ b/weed/server/qa_block_cp11b1_adversarial_test.go @@ -30,7 +30,7 @@ func qaPresetMaster(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), diff --git a/weed/server/qa_block_cp11b2_adversarial_test.go b/weed/server/qa_block_cp11b2_adversarial_test.go index 17097b669..d2dab8dbb 100644 --- a/weed/server/qa_block_cp11b2_adversarial_test.go +++ b/weed/server/qa_block_cp11b2_adversarial_test.go @@ -150,7 +150,7 @@ func TestQA_CP11B2_PlanThenCreate_OrderedCandidateParity(t *testing.T) { // Record which servers create tries, in order. var createAttempts []string - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { createAttempts = append(createAttempts, server) return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), @@ -198,7 +198,7 @@ func TestQA_CP11B2_PlanThenCreate_ReplicaOrderParity(t *testing.T) { ms := qaPlanMaster(t) var allocOrder []string - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { allocOrder = append(allocOrder, server) return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), @@ -252,7 +252,7 @@ func TestQA_CP11B2_Create_FallbackOnRPCFailure(t *testing.T) { ms := qaPlanMaster(t) callCount := 0 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount++ if callCount == 1 { return nil, fmt.Errorf("simulated RPC failure") @@ -574,7 +574,7 @@ func TestQA_CP11B2_FailedPrimary_TriedAsReplica(t *testing.T) { ms := qaPlanMaster(t) var allocLog []string callCount := 0 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { allocLog = append(allocLog, server) callCount++ if callCount == 1 { diff --git a/weed/server/qa_block_cp62_test.go b/weed/server/qa_block_cp62_test.go index 6e291c725..24128d6b7 100644 --- a/weed/server/qa_block_cp62_test.go +++ b/weed/server/qa_block_cp62_test.go @@ -319,7 +319,7 @@ func TestQA_Master_AllVSFailNoOrphan(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs2:9333") ms.blockRegistry.MarkBlockCapable("vs3:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return nil, fmt.Errorf("disk full on %s", server) } @@ -348,7 +348,7 @@ func TestQA_Master_SlowAllocateBlocksSecond(t *testing.T) { ms.blockRegistry.MarkBlockCapable("vs1:9333") var allocCount atomic.Int32 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { allocCount.Add(1) time.Sleep(100 * time.Millisecond) // simulate slow VS return &blockAllocResult{ diff --git a/weed/server/qa_block_cp63_test.go b/weed/server/qa_block_cp63_test.go index 7ae247a22..e3f9bf82b 100644 --- a/weed/server/qa_block_cp63_test.go +++ b/weed/server/qa_block_cp63_test.go @@ -24,7 +24,7 @@ func testMSForQA(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), diff --git a/weed/server/qa_block_cp82_adversarial_test.go b/weed/server/qa_block_cp82_adversarial_test.go index 91a11acfc..8d1037261 100644 --- a/weed/server/qa_block_cp82_adversarial_test.go +++ b/weed/server/qa_block_cp82_adversarial_test.go @@ -27,7 +27,7 @@ func qaCP82Master(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), diff --git a/weed/server/qa_block_cp831_adversarial_test.go b/weed/server/qa_block_cp831_adversarial_test.go index 9a1706a22..d1ccf60a1 100644 --- a/weed/server/qa_block_cp831_adversarial_test.go +++ b/weed/server/qa_block_cp831_adversarial_test.go @@ -244,7 +244,7 @@ func TestQA_CP831_SyncAll_RF3_PartialReplica_OneOfTwo_Fails(t *testing.T) { // First call succeeds (primary), second succeeds (replica 1), // third fails (replica 2). callCount := 0 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount++ if callCount == 3 { return nil, fmt.Errorf("disk full on third server") @@ -300,7 +300,7 @@ func TestQA_CP831_SyncQuorum_RF3_OneReplicaOK_Succeeds(t *testing.T) { ctx := context.Background() callCount := 0 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount++ if callCount == 3 { return nil, fmt.Errorf("disk full on third server") @@ -345,7 +345,7 @@ func TestQA_CP831_SyncQuorum_RF3_AllReplicasFail_Fails(t *testing.T) { ctx := context.Background() callCount := 0 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount++ if callCount > 1 { // primary succeeds, all replicas fail return nil, fmt.Errorf("disk full") @@ -395,7 +395,7 @@ func TestQA_CP831_ConcurrentCreate_SameName_DifferentModes(t *testing.T) { ctx := context.Background() // Slow down allocation to increase race window. - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { time.Sleep(10 * time.Millisecond) return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), @@ -593,7 +593,7 @@ func TestQA_CP831_BestEffort_NoReplicaCreate_StillSucceeds(t *testing.T) { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -627,7 +627,7 @@ func TestQA_CP831_SyncAll_SingleServer_Fails(t *testing.T) { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -668,7 +668,7 @@ func TestQA_CP831_CleanupPartialCreate_DeletesFails_NoRegistryLeak(t *testing.T) ctx := context.Background() callCount := 0 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount++ if callCount > 1 { // all replicas fail return nil, fmt.Errorf("no space") diff --git a/weed/server/qa_block_csi_lifecycle_test.go b/weed/server/qa_block_csi_lifecycle_test.go index 56882a438..652148047 100644 --- a/weed/server/qa_block_csi_lifecycle_test.go +++ b/weed/server/qa_block_csi_lifecycle_test.go @@ -104,7 +104,7 @@ func newCSILifecycleSetup(t *testing.T) (*bsi.ExportedControllerServer, *bsi.Exp ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { sanitized := strings.ReplaceAll(server, ":", "_") serverDir := filepath.Join(dir, sanitized) os.MkdirAll(serverDir, 0755) diff --git a/weed/server/qa_block_disturbance_test.go b/weed/server/qa_block_disturbance_test.go index 305b1ac25..d1e1550ca 100644 --- a/weed/server/qa_block_disturbance_test.go +++ b/weed/server/qa_block_disturbance_test.go @@ -48,7 +48,7 @@ func newDisturbanceSetup(t *testing.T) *disturbanceSetup { ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { sanitized := strings.ReplaceAll(server, ":", "_") serverDir := filepath.Join(dir, sanitized) os.MkdirAll(serverDir, 0755) diff --git a/weed/server/qa_block_durability_test.go b/weed/server/qa_block_durability_test.go index 15f5aef7d..dd1b20298 100644 --- a/weed/server/qa_block_durability_test.go +++ b/weed/server/qa_block_durability_test.go @@ -27,7 +27,7 @@ func qaDurabilityMaster(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), @@ -389,7 +389,7 @@ func TestDurability_SyncAll_PartialReplicaFails(t *testing.T) { // Make replica allocation always fail. callCount := 0 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount++ if callCount > 1 { // Replica calls fail. @@ -442,7 +442,7 @@ func TestDurability_BestEffort_PartialReplicaOK(t *testing.T) { // Make replica allocation fail. callCount := 0 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount++ if callCount > 1 { return nil, fmt.Errorf("disk full") diff --git a/weed/server/qa_block_expand_adversarial_test.go b/weed/server/qa_block_expand_adversarial_test.go index e87ce67a6..0c1a1bde7 100644 --- a/weed/server/qa_block_expand_adversarial_test.go +++ b/weed/server/qa_block_expand_adversarial_test.go @@ -28,7 +28,7 @@ func qaExpandMaster(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), diff --git a/weed/server/qa_block_nvme_publication_test.go b/weed/server/qa_block_nvme_publication_test.go index de4e3ac8a..a8b5ae1ed 100644 --- a/weed/server/qa_block_nvme_publication_test.go +++ b/weed/server/qa_block_nvme_publication_test.go @@ -455,7 +455,7 @@ func nvmeIntegrationMaster(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { // Simulate volume servers with NVMe enabled. // Each server has NVMe on :4420 and a deterministic NQN. host := server[:strings.Index(server, ":")] @@ -680,7 +680,7 @@ func TestIntegration_NVMe_MixedCluster(t *testing.T) { blockFailover: newBlockFailoverState(), } callCount := 0 - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { callCount++ host := server[:strings.Index(server, ":")] result := &blockAllocResult{ diff --git a/weed/server/qa_block_publication_test.go b/weed/server/qa_block_publication_test.go index c65cebdcb..3cd14aec8 100644 --- a/weed/server/qa_block_publication_test.go +++ b/weed/server/qa_block_publication_test.go @@ -40,7 +40,7 @@ func newPublicationMaster(t *testing.T, nvmeEnabled bool) *MasterServer { ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { sanitized := strings.ReplaceAll(server, ":", "_") serverDir := filepath.Join(dir, sanitized) os.MkdirAll(serverDir, 0755) diff --git a/weed/server/qa_block_restore_test.go b/weed/server/qa_block_restore_test.go index 6eb6457a0..1aa5e2805 100644 --- a/weed/server/qa_block_restore_test.go +++ b/weed/server/qa_block_restore_test.go @@ -42,7 +42,7 @@ func newRestoreMaster(t *testing.T) (*MasterServer, *storage.BlockVolumeStore, * ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { sanitized := strings.ReplaceAll(server, ":", "_") serverDir := filepath.Join(dir, sanitized) os.MkdirAll(serverDir, 0755) diff --git a/weed/server/qa_block_rf3_test.go b/weed/server/qa_block_rf3_test.go index 960af4b8b..24560af03 100644 --- a/weed/server/qa_block_rf3_test.go +++ b/weed/server/qa_block_rf3_test.go @@ -25,7 +25,7 @@ func qaRF3Master(t *testing.T) *MasterServer { blockAssignmentQueue: NewBlockAssignmentQueue(), blockFailover: newBlockFailoverState(), } - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { return &blockAllocResult{ Path: fmt.Sprintf("/data/%s.blk", name), IQN: fmt.Sprintf("iqn.2024.test:%s", name), diff --git a/weed/server/qa_block_snapshot_product_test.go b/weed/server/qa_block_snapshot_product_test.go index 3f5bc69fb..ef8a50b2b 100644 --- a/weed/server/qa_block_snapshot_product_test.go +++ b/weed/server/qa_block_snapshot_product_test.go @@ -61,7 +61,7 @@ func newSnapshotTestSetup(t *testing.T) *snapshotTestSetup { s := &snapshotTestSetup{ms: ms, bs: bs, store: store, dir: dir} // Wire master allocator to create real volumes. - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { sanitized := strings.ReplaceAll(server, ":", "_") serverDir := filepath.Join(dir, sanitized) if err := os.MkdirAll(serverDir, 0755); err != nil { diff --git a/weed/server/qa_block_soak_test.go b/weed/server/qa_block_soak_test.go index 7d685f386..e9afa79c4 100644 --- a/weed/server/qa_block_soak_test.go +++ b/weed/server/qa_block_soak_test.go @@ -48,7 +48,7 @@ func newSoakSetup(t *testing.T) *soakSetup { ms.blockRegistry.MarkBlockCapable("vs1:9333") ms.blockRegistry.MarkBlockCapable("vs2:9333") - ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { + ms.blockVSAllocate = func(ctx context.Context, server string, name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (*blockAllocResult, error) { sanitized := strings.ReplaceAll(server, ":", "_") serverDir := filepath.Join(dir, sanitized) os.MkdirAll(serverDir, 0755) diff --git a/weed/server/volume_grpc_block.go b/weed/server/volume_grpc_block.go index 5d069bbbd..c28cc6431 100644 --- a/weed/server/volume_grpc_block.go +++ b/weed/server/volume_grpc_block.go @@ -21,7 +21,7 @@ func (vs *VolumeServer) AllocateBlockVolume(_ context.Context, req *volume_serve return nil, fmt.Errorf("size_bytes must be > 0") } - path, iqn, iscsiAddr, err := vs.blockService.CreateBlockVol(req.Name, req.SizeBytes, req.DiskType, req.DurabilityMode) + path, iqn, iscsiAddr, err := vs.blockService.CreateBlockVolWithOptions(req.Name, req.SizeBytes, req.WalSizeBytes, req.DiskType, req.DurabilityMode) if err != nil { return nil, fmt.Errorf("create block volume %q: %w", req.Name, err) } diff --git a/weed/server/volume_grpc_block_test.go b/weed/server/volume_grpc_block_test.go index d5e6bb390..60f6f901b 100644 --- a/weed/server/volume_grpc_block_test.go +++ b/weed/server/volume_grpc_block_test.go @@ -1,10 +1,14 @@ package weed_server import ( + "context" "os" "path/filepath" "strings" "testing" + + "github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb" + "github.com/seaweedfs/seaweedfs/weed/storage/blockvol" ) func newTestBlockServiceWithDir(t *testing.T) (*BlockService, string) { @@ -46,6 +50,31 @@ func TestVS_AllocateBlockVolume(t *testing.T) { } } +func TestVS_AllocateBlockVolume_WithWalSize(t *testing.T) { + bs, _ := newTestBlockServiceWithDir(t) + vs := &VolumeServer{blockService: bs} + + resp, err := vs.AllocateBlockVolume(context.Background(), &volume_server_pb.AllocateBlockVolumeRequest{ + Name: "test-vol-wal", + SizeBytes: 4 * 1024 * 1024, + WalSizeBytes: 8 * 1024 * 1024, + DiskType: "ssd", + }) + if err != nil { + t.Fatalf("AllocateBlockVolume: %v", err) + } + + vol, err := blockvol.OpenBlockVol(resp.Path) + if err != nil { + t.Fatalf("OpenBlockVol: %v", err) + } + defer vol.Close() + + if got := vol.Info().WALSize; got != 8*1024*1024 { + t.Fatalf("wal_size=%d, want %d", got, 8*1024*1024) + } +} + func TestVS_AllocateIdempotent(t *testing.T) { bs, _ := newTestBlockServiceWithDir(t) diff --git a/weed/server/volume_server_block.go b/weed/server/volume_server_block.go index 9981ad820..5762d8d2f 100644 --- a/weed/server/volume_server_block.go +++ b/weed/server/volume_server_block.go @@ -111,6 +111,8 @@ type BlockService struct { // TestHook: if set, invoked when the legacy direct rebuild starter is used. onLegacyStartRebuild func(path, rebuildAddr string, epoch uint64) + + blockStateNotifyCh chan bool } // V2Orchestrator returns the V2 engine orchestrator for inspection/testing. @@ -199,14 +201,30 @@ func (bs *BlockService) SetAdvertisedHost(host string) { bs.advertisedHost = host } -// WireStateChangeNotify sets up shipper state change callbacks on all -// registered volumes so that degradation/recovery triggers an immediate -// heartbeat via the provided channel. Non-blocking send (buffered chan 1). +// WireStateChangeNotify sets up volume state callbacks on all registered +// volumes so that shipper transitions and durable-boundary advances trigger an +// immediate heartbeat via the provided channel. Non-blocking send (buffered +// chan 1). func (bs *BlockService) WireStateChangeNotify(ch chan bool) { + bs.blockStateNotifyCh = ch bs.blockStore.IterateBlockVolumes(func(path string, vol *blockvol.BlockVol) { - vol.SetOnShipperStateChange(func(from, to blockvol.ReplicaState) { - bs.handleShipperStateChange(path, from, to, ch) - }) + bs.attachVolumeStateCallbacks(path, vol) + }) +} + +func (bs *BlockService) attachVolumeStateCallbacks(path string, vol *blockvol.BlockVol) { + if bs == nil || vol == nil { + return + } + ch := bs.blockStateNotifyCh + vol.SetOnShipperStateChange(func(from, to blockvol.ReplicaState) { + bs.handleShipperStateChange(path, from, to, ch) + }) + vol.SetOnBarrierAccepted(func(flushedLSN uint64) { + bs.handleBarrierAccepted(path, flushedLSN, ch) + }) + vol.SetOnBarrierRejected(func(reason string) { + bs.handleBarrierRejected(path, reason, ch) }) } @@ -217,14 +235,58 @@ func (bs *BlockService) handleShipperStateChange(path string, from, to blockvol. default: // already pending } } - if bs == nil || bs.v2Core == nil || to != blockvol.ReplicaInSync { + glog.V(0).Infof("block service: shipper state change path=%s from=%s to=%s", path, from, to) + if bs == nil || bs.v2Core == nil { return } proj, ok := bs.CoreProjection(path) if !ok || proj.Role != engine.RolePrimary { return } - bs.applyCoreEvent(engine.ShipperConnectedObserved{ID: path}) + if to == blockvol.ReplicaInSync { + bs.applyCoreEvent(engine.ShipperConnectedObserved{ID: path}) + } +} + +func (bs *BlockService) handleBarrierAccepted(path string, flushedLSN uint64, ch chan bool) { + if ch != nil { + select { + case ch <- true: + default: // already pending + } + } + if bs == nil || bs.v2Core == nil || flushedLSN == 0 { + return + } + proj, ok := bs.CoreProjection(path) + if !ok || proj.Role != engine.RolePrimary || proj.Boundary.DurableLSN >= flushedLSN { + return + } + if !proj.Readiness.ShipperConnected && bs.isPrimaryShipperConnected(path) { + bs.applyCoreEvent(engine.ShipperConnectedObserved{ID: path}) + proj, ok = bs.CoreProjection(path) + if !ok || proj.Role != engine.RolePrimary || proj.Boundary.DurableLSN >= flushedLSN { + return + } + } + bs.applyCoreEvent(engine.BarrierAccepted{ID: path, FlushedLSN: flushedLSN}) +} + +func (bs *BlockService) handleBarrierRejected(path string, reason string, ch chan bool) { + if ch != nil { + select { + case ch <- true: + default: // already pending + } + } + if bs == nil || bs.v2Core == nil || reason == "" { + return + } + proj, ok := bs.CoreProjection(path) + if !ok || proj.Role != engine.RolePrimary { + return + } + bs.applyCoreEvent(engine.BarrierRejected{ID: path, Reason: reason}) } // StartBlockService scans blockDir for .blk files, opens them as block volumes, @@ -397,6 +459,12 @@ func (bs *BlockService) NQN(name string) string { // and iSCSI TargetServer. Returns path, IQN, iSCSI addr. // Idempotent: if volume already exists with same or larger size, returns existing info. func (bs *BlockService) CreateBlockVol(name string, sizeBytes uint64, diskType string, durabilityMode string) (path, iqn, iscsiAddr string, err error) { + return bs.CreateBlockVolWithOptions(name, sizeBytes, 0, diskType, durabilityMode) +} + +// CreateBlockVolWithOptions creates a new .blk file with the requested geometry, +// including an optional WAL size override. +func (bs *BlockService) CreateBlockVolWithOptions(name string, sizeBytes uint64, walSizeBytes uint64, diskType string, durabilityMode string) (path, iqn, iscsiAddr string, err error) { sanitized := blockvol.SanitizeFilename(name) path = filepath.Join(bs.blockDir, sanitized+".blk") iqn = bs.iqnPrefix + blockvol.SanitizeIQN(name) @@ -404,6 +472,7 @@ func (bs *BlockService) CreateBlockVol(name string, sizeBytes uint64, diskType s // Check if already registered. if vol, ok := bs.blockStore.GetBlockVolume(path); ok { + bs.attachVolumeStateCallbacks(path, vol) info := vol.Info() if info.VolumeSize < sizeBytes { return "", "", "", fmt.Errorf("block volume %q exists with size %d (requested %d)", @@ -437,6 +506,7 @@ func (bs *BlockService) CreateBlockVol(name string, sizeBytes uint64, diskType s } created, err := blockvol.CreateBlockVol(path, blockvol.CreateOptions{ VolumeSize: sizeBytes, + WALSize: walSizeBytes, DurabilityMode: durMode, }) if err != nil { @@ -450,6 +520,7 @@ func (bs *BlockService) CreateBlockVol(name string, sizeBytes uint64, diskType s os.Remove(path) return "", "", "", fmt.Errorf("register block volume: %w", err) } + bs.attachVolumeStateCallbacks(path, vol) adapter := blockvol.NewBlockVolAdapter(vol) bs.targetServer.AddVolume(iqn, adapter) @@ -781,12 +852,13 @@ func (bs *BlockService) applyCoreEvent(ev engine.Event) { // so the VS log contains a complete trace for post-run diagnosis. func (bs *BlockService) coreApplyAndLog(ev engine.Event) engine.ApplyResult { result := bs.v2Core.ApplyEvent(ev) - glog.V(0).Infof("core [%s]: event=%T mode=%s pub=%v reason=%q readiness={applied=%v shipper_cfg=%v shipper_conn=%v recv=%v} boundary={durable=%d committed=%d} cmds=%d", + glog.V(0).Infof("core [%s]: event=%T mode=%s pub=%v reason=%q readiness={applied=%v shipper_cfg=%v shipper_conn=%v recv=%v} boundary={durable=%d committed=%d last_barrier_ok=%v last_barrier_reason=%q} cmds=%d", ev.VolumeID(), ev, result.Projection.Mode.Name, result.Projection.Publication.Healthy, result.Projection.Publication.Reason, result.Projection.Readiness.RoleApplied, result.Projection.Readiness.ShipperConfigured, result.Projection.Readiness.ShipperConnected, result.Projection.Readiness.ReceiverReady, result.Projection.Boundary.DurableLSN, result.Projection.Boundary.CommittedLSN, + result.Projection.Boundary.LastBarrierOK, result.Projection.Boundary.LastBarrierReason, len(result.Commands)) return result } diff --git a/weed/server/volume_server_block_test.go b/weed/server/volume_server_block_test.go index 30be86e7a..0d698c3c5 100644 --- a/weed/server/volume_server_block_test.go +++ b/weed/server/volume_server_block_test.go @@ -1,6 +1,7 @@ package weed_server import ( + "bytes" "path/filepath" "reflect" "strings" @@ -1710,6 +1711,209 @@ func TestBlockService_ShipperStateChange_InSyncEmitsCoreConnectedObservation(t * } } +func TestBlockService_BarrierRejectedCallback_UpdatesCoreProjection(t *testing.T) { + bs := newTestBlockServiceDirect(t) + path := createTestVolDirect(t, bs, "vol-barrier-rejected-callback") + ch := make(chan bool, 1) + bs.WireStateChangeNotify(ch) + + errs := bs.ApplyAssignments([]blockvol.BlockVolumeAssignment{ + { + Path: path, + Epoch: 1, + Role: blockvol.RoleToWire(blockvol.RolePrimary), + LeaseTtlMs: 30000, + ReplicaServerID: "vs-2", + ReplicaDataAddr: "10.0.0.2:4260", + ReplicaCtrlAddr: "10.0.0.2:4261", + }, + }) + if len(errs) != 1 || errs[0] != nil { + t.Fatalf("apply assignment errs=%v", errs) + } + + bs.handleBarrierRejected(path, "barrier_timeout", ch) + + select { + case <-ch: + default: + t.Fatal("expected immediate heartbeat notification") + } + + after, ok := bs.CoreProjection(path) + if !ok { + t.Fatal("expected core projection after barrier rejection") + } + if after.Mode.Name != engine.ModeDegraded { + t.Fatalf("mode=%s, want %s", after.Mode.Name, engine.ModeDegraded) + } + if after.Publication.Reason != "barrier_timeout" { + t.Fatalf("reason=%q, want %q", after.Publication.Reason, "barrier_timeout") + } + if after.Boundary.LastBarrierOK { + t.Fatalf("last_barrier_ok=%v, want false", after.Boundary.LastBarrierOK) + } + if after.Boundary.LastBarrierReason != "barrier_timeout" { + t.Fatalf("last_barrier_reason=%q, want %q", after.Boundary.LastBarrierReason, "barrier_timeout") + } +} + +func TestBlockService_CreateBlockVol_WiresStateChangeCallbackForNewVolumes(t *testing.T) { + dir := t.TempDir() + bs := StartBlockService("127.0.0.1:0", dir, "iqn.2024-01.com.test:vol.", "127.0.0.1:3260,1", NVMeConfig{}) + if bs == nil { + t.Fatal("expected non-nil BlockService") + } + defer bs.Shutdown() + + ch := make(chan bool, 1) + bs.WireStateChangeNotify(ch) + + path, _, _, err := bs.CreateBlockVol("vol-new-callback", 4*1024*1024, "", "sync_all") + if err != nil { + t.Fatalf("CreateBlockVol: %v", err) + } + + primary, ok := bs.blockStore.GetBlockVolume(path) + if !ok { + t.Fatalf("created volume %s not found", path) + } + + replicaPath := filepath.Join(t.TempDir(), "replica.blockvol") + replica, err := blockvol.CreateBlockVol(replicaPath, blockvol.CreateOptions{ + VolumeSize: 4 * 1024 * 1024, + DurabilityMode: blockvol.DurabilitySyncAll, + }) + if err != nil { + t.Fatalf("CreateBlockVol replica: %v", err) + } + defer replica.Close() + replica.SetRole(blockvol.RoleReplica) + if err := replica.SetEpoch(1); err != nil { + t.Fatalf("replica SetEpoch: %v", err) + } + replica.SetMasterEpoch(1) + + if err := replica.StartReplicaReceiver("127.0.0.1:0", "127.0.0.1:0"); err != nil { + t.Fatalf("StartReplicaReceiver: %v", err) + } + recv := replica.ReplicaReceiverAddr() + if recv == nil { + t.Fatal("ReplicaReceiverAddr returned nil") + } + + errs := bs.ApplyAssignments([]blockvol.BlockVolumeAssignment{{ + Path: path, + Epoch: 1, + Role: blockvol.RoleToWire(blockvol.RolePrimary), + LeaseTtlMs: 30000, + ReplicaServerID: "vs-2", + ReplicaDataAddr: recv.DataAddr, + ReplicaCtrlAddr: recv.CtrlAddr, + }}) + if len(errs) != 1 || errs[0] != nil { + t.Fatalf("ApplyAssignments errs=%v", errs) + } + + if err := primary.WriteLBA(0, bytes.Repeat([]byte{'N'}, 4096)); err != nil { + t.Fatalf("WriteLBA: %v", err) + } + if err := primary.SyncCache(); err != nil { + t.Fatalf("SyncCache: %v", err) + } + + select { + case <-ch: + case <-time.After(2 * time.Second): + t.Fatal("expected immediate heartbeat notification for newly created volume") + } +} + +func TestBlockService_CreateBlockVol_SyncCacheSuccessEmitsBarrierAccepted(t *testing.T) { + dir := t.TempDir() + bs := StartBlockService("127.0.0.1:0", dir, "iqn.2024-01.com.test:vol.", "127.0.0.1:3260,1", NVMeConfig{}) + if bs == nil { + t.Fatal("expected non-nil BlockService") + } + defer bs.Shutdown() + + ch := make(chan bool, 1) + bs.WireStateChangeNotify(ch) + + path, _, _, err := bs.CreateBlockVol("vol-new-barrier-callback", 4*1024*1024, "", "sync_all") + if err != nil { + t.Fatalf("CreateBlockVol: %v", err) + } + + primary, ok := bs.blockStore.GetBlockVolume(path) + if !ok { + t.Fatalf("created volume %s not found", path) + } + + replicaPath := filepath.Join(t.TempDir(), "replica-barrier.blockvol") + replica, err := blockvol.CreateBlockVol(replicaPath, blockvol.CreateOptions{ + VolumeSize: 4 * 1024 * 1024, + DurabilityMode: blockvol.DurabilitySyncAll, + }) + if err != nil { + t.Fatalf("CreateBlockVol replica: %v", err) + } + defer replica.Close() + replica.SetRole(blockvol.RoleReplica) + if err := replica.SetEpoch(1); err != nil { + t.Fatalf("replica SetEpoch: %v", err) + } + replica.SetMasterEpoch(1) + + if err := replica.StartReplicaReceiver("127.0.0.1:0", "127.0.0.1:0"); err != nil { + t.Fatalf("StartReplicaReceiver: %v", err) + } + recv := replica.ReplicaReceiverAddr() + if recv == nil { + t.Fatal("ReplicaReceiverAddr returned nil") + } + + errs := bs.ApplyAssignments([]blockvol.BlockVolumeAssignment{{ + Path: path, + Epoch: 1, + Role: blockvol.RoleToWire(blockvol.RolePrimary), + LeaseTtlMs: 30000, + ReplicaServerID: "vs-2", + ReplicaDataAddr: recv.DataAddr, + ReplicaCtrlAddr: recv.CtrlAddr, + }}) + if len(errs) != 1 || errs[0] != nil { + t.Fatalf("ApplyAssignments errs=%v", errs) + } + + if err := primary.WriteLBA(0, bytes.Repeat([]byte{'B'}, 4096)); err != nil { + t.Fatalf("WriteLBA: %v", err) + } + if err := primary.SyncCache(); err != nil { + t.Fatalf("SyncCache: %v", err) + } + + deadline := time.Now().Add(2 * time.Second) + for time.Now().Before(deadline) { + proj, ok := bs.CoreProjection(path) + if ok && proj.Boundary.DurableLSN > 0 && proj.Publication.Healthy { + return + } + time.Sleep(10 * time.Millisecond) + } + + proj, ok := bs.CoreProjection(path) + if !ok { + t.Fatal("expected core projection after SyncCache success") + } + if proj.Boundary.DurableLSN == 0 { + t.Fatalf("durable_lsn=%d after SyncCache success", proj.Boundary.DurableLSN) + } + if !proj.Publication.Healthy { + t.Fatalf("expected publish_healthy after SyncCache success, projection=%+v", proj) + } +} + func TestBlockService_ObservePrimaryShipperConnectivityStatus_EmitsCoreConnectedObservation(t *testing.T) { bs := newTestBlockServiceDirect(t) path := createTestVolDirect(t, bs, "vol-shipper-connected-recheck") diff --git a/weed/storage/blockvol/blockapi/client_test.go b/weed/storage/blockvol/blockapi/client_test.go index 970da58c1..f48d23309 100644 --- a/weed/storage/blockvol/blockapi/client_test.go +++ b/weed/storage/blockvol/blockapi/client_test.go @@ -24,6 +24,9 @@ func TestClientCreateVolume(t *testing.T) { if req.Name != "test-vol" { t.Errorf("expected name test-vol, got %s", req.Name) } + if req.WALSizeBytes != 256<<20 { + t.Errorf("expected wal_size_bytes %d, got %d", 256<<20, req.WALSizeBytes) + } w.Header().Set("Content-Type", "application/json") json.NewEncoder(w).Encode(VolumeInfo{ Name: req.Name, @@ -38,8 +41,9 @@ func TestClientCreateVolume(t *testing.T) { client := NewClient(ts.URL) info, err := client.CreateVolume(context.Background(), CreateVolumeRequest{ - Name: "test-vol", - SizeBytes: 1 << 30, + Name: "test-vol", + SizeBytes: 1 << 30, + WALSizeBytes: 256 << 20, }) if err != nil { t.Fatal(err) diff --git a/weed/storage/blockvol/blockapi/types.go b/weed/storage/blockvol/blockapi/types.go index fcd705bf2..9ac74ce55 100644 --- a/weed/storage/blockvol/blockapi/types.go +++ b/weed/storage/blockvol/blockapi/types.go @@ -10,6 +10,7 @@ import ( type CreateVolumeRequest struct { Name string `json:"name"` SizeBytes uint64 `json:"size_bytes"` + WALSizeBytes uint64 `json:"wal_size_bytes,omitempty"` ReplicaPlacement string `json:"replica_placement"` // SeaweedFS placement string: "000", "001", "010", "100" DiskType string `json:"disk_type"` // e.g. "ssd", "hdd" DurabilityMode string `json:"durability_mode,omitempty"` // "best_effort", "sync_all", "sync_quorum" @@ -46,10 +47,10 @@ type VolumeInfo struct { // CP11B-4: Operator-facing health state. HealthState string `json:"health_state"` // "healthy", "degraded", "rebuilding", "unsafe" // CP13-9: Normalized volume mode for constrained-runtime surfaces. - VolumeMode string `json:"volume_mode,omitempty"` // "allocated_only", "bootstrap_pending", "publish_healthy", "degraded", "needs_rebuild" - VolumeModeReason string `json:"volume_mode_reason,omitempty"` - EngineProjectionMode string `json:"engine_projection_mode,omitempty"` // T1: VS-local V2 engine projection - ClusterReplicationMode string `json:"cluster_replication_mode,omitempty"` // T5: master-owned cluster RF2 health + VolumeMode string `json:"volume_mode,omitempty"` // "allocated_only", "bootstrap_pending", "publish_healthy", "degraded", "needs_rebuild" + VolumeModeReason string `json:"volume_mode_reason,omitempty"` + EngineProjectionMode string `json:"engine_projection_mode,omitempty"` // T1: VS-local V2 engine projection + ClusterReplicationMode string `json:"cluster_replication_mode,omitempty"` // T5: master-owned cluster RF2 health } // ResolvedPolicyResponse is the response for POST /block/volume/resolve. diff --git a/weed/storage/blockvol/blockvol.go b/weed/storage/blockvol/blockvol.go index 8cf1e6c79..73beb43c7 100644 --- a/weed/storage/blockvol/blockvol.go +++ b/weed/storage/blockvol/blockvol.go @@ -92,6 +92,14 @@ type BlockVol struct { // Shipper state change callback — triggers immediate heartbeat. onShipperStateChange func(from, to ReplicaState) + // Barrier acceptance callback — reports authoritative durable progress after + // a successful SyncCache/group-commit fence. + onBarrierAccepted func(flushedLSN uint64) + + // Barrier rejection callback — reports the semantic reason for a failed + // distributed durability fence. + onBarrierRejected func(reason string) + // liveShippingPolicy gates whether configured shippers may consume current // live-tail WAL entries. The host uses this to keep replicas in bounded // catch-up until their active protocol session reaches a live-eligible phase. @@ -208,7 +216,7 @@ func CreateBlockVol(path string, opts CreateOptions, cfgs ...BlockVolConfig) (*B if v.shipperGroup != nil { v.shipperGroup.EvaluateRetentionBudgets(RetentionBudgetParams{ Timeout: walRetentionTimeout, - MaxBytes: walRetentionMaxBytes, + MaxBytes: 0, // CP13-6 max-bytes disabled: uses replicaFlushedLSN which can't advance without barrier; v2 will replace with negotiated recovery protocol PrimaryHeadLSN: v.nextLSN.Load() - 1, BlockSize: v.super.BlockSize, }) @@ -336,7 +344,7 @@ func OpenBlockVol(path string, cfgs ...BlockVolConfig) (*BlockVol, error) { if v.shipperGroup != nil { v.shipperGroup.EvaluateRetentionBudgets(RetentionBudgetParams{ Timeout: walRetentionTimeout, - MaxBytes: walRetentionMaxBytes, + MaxBytes: 0, // CP13-6 max-bytes disabled: uses replicaFlushedLSN which can't advance without barrier; v2 will replace with negotiated recovery protocol PrimaryHeadLSN: v.nextLSN.Load() - 1, BlockSize: v.super.BlockSize, }) @@ -820,7 +828,15 @@ func (v *BlockVol) SyncCache() error { defer v.endOp() v.ioMu.RLock() defer v.ioMu.RUnlock() - return v.groupCommit.Submit() + if err := v.groupCommit.Submit(); err != nil { + return err + } + if v.onBarrierAccepted != nil { + if flushedLSN, ok := v.currentDurableBoundary(); ok && flushedLSN > 0 { + v.onBarrierAccepted(flushedLSN) + } + } + return nil } // ReplicaAddr holds the data and control addresses for one replica. @@ -874,6 +890,38 @@ func (a *walAccess) StreamEntries(fromLSN uint64, fn func(*WALEntry) error) erro // Called by the volume server to trigger immediate heartbeat on degradation/recovery. func (v *BlockVol) SetOnShipperStateChange(fn func(from, to ReplicaState)) { v.onShipperStateChange = fn + if v != nil && v.shipperGroup != nil { + v.shipperGroup.SetOnStateChange(fn) + } +} + +// SetOnBarrierAccepted registers a callback invoked after SyncCache returns +// success with a positive durable boundary. +func (v *BlockVol) SetOnBarrierAccepted(fn func(flushedLSN uint64)) { + v.onBarrierAccepted = fn +} + +// SetOnBarrierRejected registers a callback invoked when a distributed barrier +// fails and reports a stable semantic reason. +func (v *BlockVol) SetOnBarrierRejected(fn func(reason string)) { + v.onBarrierRejected = fn + if v != nil && v.shipperGroup != nil { + v.shipperGroup.SetOnBarrierFailure(fn) + } +} + +func (v *BlockVol) currentDurableBoundary() (uint64, bool) { + if v == nil { + return 0, false + } + if v.DurabilityMode() == DurabilitySyncAll && v.shipperGroup != nil && v.shipperGroup.Len() > 0 { + return v.shipperGroup.MinReplicaFlushedLSNAll() + } + headLSN := v.nextLSN.Load() + if headLSN == 0 { + return 0, false + } + return headLSN - 1, true } // SetLiveShippingPolicy installs a host-provided gate for current live-tail @@ -949,6 +997,9 @@ func (v *BlockVol) SetReplicaAddrs(addrs []ReplicaAddr) { if v.onShipperStateChange != nil { v.shipperGroup.SetOnStateChange(v.onShipperStateChange) } + if v.onBarrierRejected != nil { + v.shipperGroup.SetOnBarrierFailure(v.onBarrierRejected) + } if v.liveShippingPolicy != nil { v.shipperGroup.SetLiveShippingPolicy(v.liveShippingPolicy) } diff --git a/weed/storage/blockvol/shipper_group.go b/weed/storage/blockvol/shipper_group.go index d6ecf8c4e..7ffa773b4 100644 --- a/weed/storage/blockvol/shipper_group.go +++ b/weed/storage/blockvol/shipper_group.go @@ -297,6 +297,16 @@ func (sg *ShipperGroup) SetOnStateChange(fn func(from, to ReplicaState)) { } } +// SetOnBarrierFailure registers a callback on all current shippers for failed +// barrier attempts. +func (sg *ShipperGroup) SetOnBarrierFailure(fn func(reason string)) { + sg.mu.RLock() + defer sg.mu.RUnlock() + for _, s := range sg.shippers { + s.SetOnBarrierFailure(fn) + } +} + // SetLiveShippingPolicy installs a host-provided gate on all current shippers. // The policy is evaluated before a live-tail WAL entry is dialed or sent. func (sg *ShipperGroup) SetLiveShippingPolicy(fn func(replicaID string, entryLSN uint64) (allow bool, reason string)) { diff --git a/weed/storage/blockvol/sync_all_bug_test.go b/weed/storage/blockvol/sync_all_bug_test.go index 8045df1a5..2ed333449 100644 --- a/weed/storage/blockvol/sync_all_bug_test.go +++ b/weed/storage/blockvol/sync_all_bug_test.go @@ -371,6 +371,59 @@ func TestSyncAll_MultipleFlush_NoWritesBetween(t *testing.T) { } } +// TestSyncAll_LateConfiguredShipper_FirstBarrierReplaysBacklog captures the +// integrated Stage 0 shape: writes can land before the VS finishes shipper +// configuration, and the first fsync-triggered barrier must replay that +// retained backlog instead of failing with a remote I/O error. +func TestSyncAll_LateConfiguredShipper_FirstBarrierReplaysBacklog(t *testing.T) { + primary, replica := createSyncAllPair(t) + defer primary.Close() + defer replica.Close() + + // Writes happen before any replica transport is configured. + for i := 0; i < 4; i++ { + if err := primary.WriteLBA(uint64(i), makeBlock(byte('a'+i))); err != nil { + t.Fatalf("pre-config write %d: %v", i, err) + } + } + + recv, err := NewReplicaReceiver(replica, "127.0.0.1:0", "127.0.0.1:0") + if err != nil { + t.Fatal(err) + } + recv.Serve() + defer recv.Stop() + + // Configure the shipper after backlog already exists. + primary.SetReplicaAddr(recv.DataAddr(), recv.CtrlAddr()) + + syncDone := make(chan error, 1) + go func() { + syncDone <- primary.SyncCache() + }() + + select { + case err := <-syncDone: + if err != nil { + t.Fatalf("SyncCache after late shipper config failed: %v", err) + } + case <-time.After(10 * time.Second): + t.Fatal("SyncCache after late shipper config hung") + } + + replica.flusher.FlushOnce() + for i := 0; i < 4; i++ { + got, err := replica.ReadLBA(uint64(i), 4096) + if err != nil { + t.Fatalf("replica ReadLBA(%d): %v", i, err) + } + expected := byte('a' + i) + if got[0] != expected { + t.Fatalf("replica LBA %d: expected %c, got %c", i, expected, got[0]) + } + } +} + // --- Helpers --- func createSyncAllPair(t *testing.T) (primary *BlockVol, replica *BlockVol) { diff --git a/weed/storage/blockvol/testrunner/actions/devops.go b/weed/storage/blockvol/testrunner/actions/devops.go index 95f91887c..a8ed85f2a 100644 --- a/weed/storage/blockvol/testrunner/actions/devops.go +++ b/weed/storage/blockvol/testrunner/actions/devops.go @@ -11,8 +11,8 @@ import ( "strings" "time" - "github.com/seaweedfs/seaweedfs/weed/storage/blockvol/testrunner/internal/blockapi" tr "github.com/seaweedfs/seaweedfs/weed/storage/blockvol/testrunner" + "github.com/seaweedfs/seaweedfs/weed/storage/blockvol/testrunner/internal/blockapi" ) // RegisterDevOpsActions registers SeaweedFS cluster management actions. @@ -297,7 +297,8 @@ func waitClusterReady(ctx context.Context, actx *tr.ActionContext, act tr.Action } // createBlockVolume creates a block volume via the master block API. -// Params: name, size (human e.g. "50M") or size_bytes, replica_factor (default 1). +// Params: name, size (human e.g. "50M") or size_bytes, wal_size (human) or +// wal_size_bytes, replica_factor (default 1). // Sets save_as=JSON, save_as_capacity, save_as_iscsi_addr, save_as_iqn. func createBlockVolume(ctx context.Context, actx *tr.ActionContext, act tr.Action) (map[string]string, error) { client, err := blockAPIClient(actx, act) @@ -327,6 +328,19 @@ func createBlockVolume(ctx context.Context, actx *tr.ActionContext, act tr.Actio } } + var walSizeBytes uint64 + if wsb := act.Params["wal_size_bytes"]; wsb != "" { + walSizeBytes, err = strconv.ParseUint(wsb, 10, 64) + if err != nil { + return nil, fmt.Errorf("create_block_volume: invalid wal_size_bytes: %w", err) + } + } else if ws := act.Params["wal_size"]; ws != "" { + walSizeBytes, err = ParseSizeBytes(ws) + if err != nil { + return nil, fmt.Errorf("create_block_volume: invalid wal_size: %w", err) + } + } + rf := ParseInt(act.Params["replica_factor"], 1) durMode := act.Params["durability_mode"] @@ -334,6 +348,7 @@ func createBlockVolume(ctx context.Context, actx *tr.ActionContext, act tr.Actio info, err := client.CreateVolume(ctx, blockapi.CreateVolumeRequest{ Name: name, SizeBytes: sizeBytes, + WALSizeBytes: walSizeBytes, ReplicaFactor: rf, DurabilityMode: durMode, }) diff --git a/weed/storage/blockvol/testrunner/actions/devops_test.go b/weed/storage/blockvol/testrunner/actions/devops_test.go index 1f62ad321..a6ff7f762 100644 --- a/weed/storage/blockvol/testrunner/actions/devops_test.go +++ b/weed/storage/blockvol/testrunner/actions/devops_test.go @@ -1,6 +1,10 @@ package actions import ( + "context" + "encoding/json" + "net/http" + "net/http/httptest" "sort" "strings" "testing" @@ -31,6 +35,8 @@ func TestDevOpsActions_Registration(t *testing.T) { "block_promote", "wait_volume_healthy", "discover_primary", + "collect_glog", + "collect_debug", } for _, name := range expected { @@ -47,8 +53,8 @@ func TestDevOpsActions_Tier(t *testing.T) { byTier := registry.ListByTier() devopsActions := byTier[tr.TierDevOps] - if len(devopsActions) != 17 { - t.Errorf("devops tier has %d actions, want 17", len(devopsActions)) + if len(devopsActions) != 19 { + t.Errorf("devops tier has %d actions, want 19", len(devopsActions)) } // Verify all are in devops tier. @@ -104,8 +110,8 @@ func TestAllActions_Registration(t *testing.T) { if n := len(byTier[tr.TierBlock]); n != 64 { t.Errorf("block: %d, want 64", n) } - if n := len(byTier[tr.TierDevOps]); n != 17 { - t.Errorf("devops: %d, want 17", n) + if n := len(byTier[tr.TierDevOps]); n != 19 { + t.Errorf("devops: %d, want 19", n) } if n := len(byTier[tr.TierChaos]); n != 5 { t.Errorf("chaos: %d, want 5", n) @@ -114,13 +120,65 @@ func TestAllActions_Registration(t *testing.T) { t.Errorf("k8s: %d, want 14", n) } - // Total should be 116 (115 prev + 1 recovery: measure_rebuild). + // Total should reflect the currently registered cross-tier action set. total := 0 for _, actions := range byTier { total += len(actions) } - if total != 117 { - t.Errorf("total actions: %d, want 117", total) + if total != 119 { + t.Errorf("total actions: %d, want 119", total) + } +} + +func TestCreateBlockVolume_ParsesAndForwardsWALSize(t *testing.T) { + var captured blockapi.CreateVolumeRequest + ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + if r.Method != http.MethodPost || r.URL.Path != "/block/volume" { + t.Fatalf("unexpected request: %s %s", r.Method, r.URL.Path) + } + if err := json.NewDecoder(r.Body).Decode(&captured); err != nil { + t.Fatalf("decode request: %v", err) + } + w.Header().Set("Content-Type", "application/json") + _ = json.NewEncoder(w).Encode(blockapi.VolumeInfo{ + Name: captured.Name, + VolumeServer: "vs1:9333", + SizeBytes: captured.SizeBytes, + ISCSIAddr: "127.0.0.1:3260", + IQN: "iqn.2024.test:vol", + }) + })) + defer ts.Close() + + actx := &tr.ActionContext{ + Vars: map[string]string{ + "master_url": ts.URL, + }, + Log: func(string, ...interface{}) {}, + } + act := tr.Action{ + Action: "create_block_volume", + Params: map[string]string{ + "name": "wal-sized-vol", + "size": "1G", + "wal_size": "256M", + "replica_factor": "2", + "durability_mode": "sync_all", + }, + SaveAs: "vol", + } + + if _, err := createBlockVolume(context.Background(), actx, act); err != nil { + t.Fatalf("createBlockVolume: %v", err) + } + if captured.Name != "wal-sized-vol" { + t.Fatalf("name=%q", captured.Name) + } + if captured.WALSizeBytes != 256<<20 { + t.Fatalf("wal_size_bytes=%d, want %d", captured.WALSizeBytes, 256<<20) + } + if actx.Vars["vol_iqn"] == "" || actx.Vars["vol_iscsi_addr"] == "" { + t.Fatalf("expected save_as vars to be populated, vars=%v", actx.Vars) } } @@ -216,9 +274,9 @@ func TestVolumeHealthyReady_AllowsSyncAllOnlyAfterPublishHealthy(t *testing.T) { wantReady: true, }, { - name: "nil_info_rejected", - info: nil, - wantReady: false, + name: "nil_info_rejected", + info: nil, + wantReady: false, wantReason: "missing", }, } diff --git a/weed/storage/blockvol/testrunner/actions/io.go b/weed/storage/blockvol/testrunner/actions/io.go index 30bf1b98b..d2f7e292f 100644 --- a/weed/storage/blockvol/testrunner/actions/io.go +++ b/weed/storage/blockvol/testrunner/actions/io.go @@ -21,6 +21,19 @@ func RegisterIOActions(r *tr.Registry) { r.RegisterFunc("stop_bg", tr.TierBlock, stopBg) } +func ddSyncConv(mode string) (string, error) { + switch mode { + case "", "fsync": + return "fsync", nil + case "fdatasync": + return "fdatasync", nil + case "none": + return "", nil + default: + return "", fmt.Errorf("unsupported sync_mode %q", mode) + } +} + // ddWrite writes random data using dd, returns the md5 checksum. func ddWrite(ctx context.Context, actx *tr.ActionContext, act tr.Action) (map[string]string, error) { device := act.Params["device"] @@ -39,6 +52,10 @@ func ddWrite(ctx context.Context, actx *tr.ActionContext, act tr.Action) (map[st if oflag == "" { oflag = "direct" } + syncConv, err := ddSyncConv(act.Params["sync_mode"]) + if err != nil { + return nil, fmt.Errorf("dd_write: %w", err) + } node, err := GetNode(actx, act.Node) if err != nil { @@ -50,20 +67,30 @@ func ddWrite(ctx context.Context, actx *tr.ActionContext, act tr.Action) (map[st if err := ensureTempRoot(ctx, node, actx); err != nil { return nil, fmt.Errorf("dd_write: %w", err) } - genCmd := fmt.Sprintf("dd if=/dev/urandom of=%s bs=%s count=%s 2>/dev/null", tmpFile, bs, count) + genCmd := fmt.Sprintf("dd if=/dev/urandom of=%s bs=%s count=%s status=none", tmpFile, bs, count) _, stderr, code, err := node.RunRoot(ctx, genCmd) if err != nil || code != 0 { return nil, fmt.Errorf("dd_write gen: code=%d stderr=%s err=%v", code, stderr, err) } - writeCmd := fmt.Sprintf("dd if=%s of=%s bs=%s oflag=%s conv=fsync", tmpFile, device, bs, oflag) + writeCmd := fmt.Sprintf("dd if=%s of=%s bs=%s oflag=%s", tmpFile, device, bs, oflag) + if syncConv != "" { + writeCmd += fmt.Sprintf(" conv=%s", syncConv) + } if seek := act.Params["seek"]; seek != "" { writeCmd += fmt.Sprintf(" seek=%s", seek) } - writeCmd += " 2>/dev/null" + writeCmd += " status=none" + if actx.Log != nil { + if syncConv == "" { + actx.Log(" dd_write: device=%s bs=%s count=%s oflag=%s sync_mode=none", device, bs, count, oflag) + } else { + actx.Log(" dd_write: device=%s bs=%s count=%s oflag=%s sync_mode=%s", device, bs, count, oflag, syncConv) + } + } _, stderr, code, err = node.RunRoot(ctx, writeCmd) if err != nil || code != 0 { - return nil, fmt.Errorf("dd_write: code=%d stderr=%s err=%v", code, stderr, err) + return nil, fmt.Errorf("dd_write: code=%d sync_mode=%s stderr=%s err=%v", code, syncConvOrNone(syncConv), stderr, err) } md5Cmd := fmt.Sprintf("md5sum %s | cut -d' ' -f1", tmpFile) @@ -105,11 +132,10 @@ func ddReadMD5(ctx context.Context, actx *tr.ActionContext, act tr.Action) (map[ if err := ensureTempRoot(ctx, node, actx); err != nil { return nil, fmt.Errorf("dd_read_md5: %w", err) } - readCmd := fmt.Sprintf("dd if=%s of=%s bs=%s count=%s iflag=direct", device, tmpFile, bs, count) + readCmd := fmt.Sprintf("dd if=%s of=%s bs=%s count=%s iflag=direct status=none", device, tmpFile, bs, count) if skip := act.Params["skip"]; skip != "" { readCmd += fmt.Sprintf(" skip=%s", skip) } - readCmd += " 2>/dev/null" _, stderr, code, err := node.RunRoot(ctx, readCmd) if err != nil || code != 0 { return nil, fmt.Errorf("dd_read_md5 read: code=%d stderr=%s err=%v", code, stderr, err) @@ -130,6 +156,13 @@ func ddReadMD5(ctx context.Context, actx *tr.ActionContext, act tr.Action) (map[ return map[string]string{"value": md5}, nil } +func syncConvOrNone(syncConv string) string { + if syncConv == "" { + return "none" + } + return syncConv +} + func fioAction(ctx context.Context, actx *tr.ActionContext, act tr.Action) (map[string]string, error) { device := act.Params["device"] if device == "" { diff --git a/weed/storage/blockvol/testrunner/cmd/sw-test-runner/main.go b/weed/storage/blockvol/testrunner/cmd/sw-test-runner/main.go index d0d6ae492..aaec4642a 100644 --- a/weed/storage/blockvol/testrunner/cmd/sw-test-runner/main.go +++ b/weed/storage/blockvol/testrunner/cmd/sw-test-runner/main.go @@ -533,6 +533,89 @@ func suiteCmd(args []string) { }) } + // --- Evidence collection --- + if saveDir != "" { + timestamp := time.Now().Format("20060102-150405") + evidenceDir := filepath.Join(saveDir, fmt.Sprintf("%s-%s", suite.Name, timestamp)) + if err := os.MkdirAll(evidenceDir, 0755); err != nil { + logger.Printf("warning: create evidence dir: %v", err) + } else { + logger.Printf("[evidence] collecting to %s", evidenceDir) + + // Pull glog from all nodes — only files created during this suite run. + glogDir := filepath.Join(evidenceDir, "glog") + os.MkdirAll(glogDir, 0755) + for nodeName, nodeRunner := range actx.Nodes { + sshNode, ok := nodeRunner.(*infra.Node) + if !ok { + continue + } + for _, pattern := range suite.Evidence.GlogPatterns { + // Only files modified in the last 10 minutes (covers the run window). + cmd := fmt.Sprintf("find /tmp -maxdepth 1 -name 'weed.*.INFO.*' -mmin -10 2>/dev/null | sort") + if !strings.Contains(pattern, "weed") { + cmd = fmt.Sprintf("ls -t %s 2>/dev/null | head -3", pattern) + } + stdout, _, _, err := nodeRunner.Run(ctx, cmd) + if err != nil { + continue + } + collected := 0 + for _, f := range strings.Split(strings.TrimSpace(stdout), "\n") { + f = strings.TrimSpace(f) + if f == "" { + continue + } + localPath := filepath.Join(glogDir, nodeName+"-"+filepath.Base(f)) + sshNode.Download(f, localPath) + collected++ + } + if collected > 0 { + logger.Printf("[evidence] %d glog files from %s", collected, nodeName) + } + } + } + + // Pull debug endpoints via SSH to the run node. + runNodeName := suite.Evidence.RunNode + if runNodeName == "" { + runNodeName = "m01" + } + if runNode, ok := actx.Nodes[runNodeName]; ok { + for _, ep := range suite.Evidence.DebugEndpoints { + stdout, _, _, err := runNode.Run(ctx, fmt.Sprintf("curl -s --max-time 3 %s 2>/dev/null", ep)) + if err != nil || len(stdout) < 2 { + continue + } + // Derive filename from endpoint URL. + epName := strings.ReplaceAll(ep, "http://", "") + epName = strings.ReplaceAll(epName, "/", "_") + epName = strings.ReplaceAll(epName, ":", "-") + localPath := filepath.Join(evidenceDir, "debug-"+epName+".json") + os.WriteFile(localPath, []byte(stdout), 0644) + logger.Printf("[evidence] debug: %s", localPath) + } + } + + // Pull result bundle from run node. + if runNode, ok := actx.Nodes[runNodeName]; ok { + if sshNode, ok := runNode.(*infra.Node); ok { + latestCmd := "ls -td results/*/ 2>/dev/null | head -1" + latestDir, _, _, _ := runNode.Run(ctx, latestCmd) + latestDir = strings.TrimSpace(latestDir) + if latestDir != "" { + for _, fname := range []string{"result.json", "result.xml", "result.html", "manifest.json", "scenario.yaml"} { + sshNode.Download(latestDir+fname, filepath.Join(evidenceDir, fname)) + } + } + } + } + + // Copy console log. + logger.Printf("[evidence] saved to %s", evidenceDir) + } + } + // --- Summary --- logger.Printf("") logger.Printf("=== SUITE RESULT: %s ===", suiteResult.Status) diff --git a/weed/storage/blockvol/testrunner/internal/blockapi/types.go b/weed/storage/blockvol/testrunner/internal/blockapi/types.go index 0b28e4c99..6bfa9a5b5 100644 --- a/weed/storage/blockvol/testrunner/internal/blockapi/types.go +++ b/weed/storage/blockvol/testrunner/internal/blockapi/types.go @@ -7,6 +7,7 @@ package blockapi type CreateVolumeRequest struct { Name string `json:"name"` SizeBytes uint64 `json:"size_bytes"` + WALSizeBytes uint64 `json:"wal_size_bytes,omitempty"` ReplicaPlacement string `json:"replica_placement"` DiskType string `json:"disk_type"` DurabilityMode string `json:"durability_mode,omitempty"` diff --git a/weed/storage/blockvol/testrunner/scenarios/internal/recovery-baseline-failover.yaml b/weed/storage/blockvol/testrunner/scenarios/internal/recovery-baseline-failover.yaml index 7f1be37c4..a081992f3 100644 --- a/weed/storage/blockvol/testrunner/scenarios/internal/recovery-baseline-failover.yaml +++ b/weed/storage/blockvol/testrunner/scenarios/internal/recovery-baseline-failover.yaml @@ -3,16 +3,22 @@ timeout: 10m # Robust dimension: automatic failover after primary death. # -# Flow: -# 1. Create RF=2 sync_all volume on the natural primary -# 2. Bootstrap first barrier with a real write, then record primary -# 3. Kill primary VS (SIGKILL) -# 4. Wait for lease expiry (30s TTL + margin) -# 5. Verify: master auto-promotes replica to primary (no manual promote) -# 6. Reconnect iSCSI to new primary, verify I/O works +# Stage 1 runtime scenario: +# 1. Create RF=2 sync_all volume +# 2. Bootstrap: small write + explicit fsync = first durability fence +# 3. wait_volume_healthy (now valid: barrier established durable truth) +# 4. Sustained workload: fio + large dd_write with checksum +# 5. Kill primary VS (SIGKILL) +# 6. Wait for lease expiry (30s TTL + margin) +# 7. Verify: master auto-promotes replica to primary +# 8. Reconnect iSCSI to new primary, verify data continuity + I/O # -# This tests the master's automatic failover path via -# evaluatePromotionLocked() in master_block_failover.go. +# Stage 0 closure now lives in `recovery-bootstrap-closure.yaml`. This file keeps +# the heavier workload and failover path together as the next-stage scenario. +# +# Key contract rule: publish_healthy requires DurableLSN > 0, which +# requires at least one successful barrier. So wait_volume_healthy +# must come AFTER the first durability fence, not before any writes. env: master_url: "http://10.0.0.3:9433" @@ -90,7 +96,10 @@ phases: replica_factor: "2" durability_mode: "sync_all" - - name: record-before + # Phase 1: Bootstrap — establish first durability fence. + # A fresh sync_all volume needs one successful barrier before + # publish_healthy can be reached (DurableLSN > 0 gate). + - name: bootstrap-fence actions: - action: discover_primary name: "{{ volume_name }}" @@ -99,13 +108,15 @@ phases: - action: print msg: "Before: primary={{ before }} ({{ before_server }}), replica={{ before_replica_node }}" - # Bootstrap sync_all with a real write before requiring publish_healthy. - # Fresh RF=2 volumes do not become publish_healthy until the first - # barrier succeeds and establishes durable truth. - action: lookup_block_volume name: "{{ volume_name }}" save_as: vol + # Wait for shipper to be configured (assignment delivered + shipper wired). + # 10s is enough for heartbeat cycle + assignment delivery + shipper setup. + - action: sleep + duration: 10s + - action: iscsi_login_direct node: m01 host: "{{ vol_iscsi_host }}" @@ -113,6 +124,31 @@ phases: iqn: "{{ vol_iqn }}" save_as: device + # Bootstrap write: small, explicit fsync = first durability fence. + # This is NOT the main workload — it's the minimal write needed to + # establish barrier truth so publish_healthy becomes reachable. + - action: dd_write + node: m01 + device: "{{ device }}" + bs: 4k + count: "1" + sync_mode: fsync + save_as: bootstrap_md5 + + - action: print + msg: "Bootstrap fence passed — first barrier confirmed (md5={{ bootstrap_md5 }})" + + # NOW wait_volume_healthy is valid: barrier has established DurableLSN > 0. + - action: wait_volume_healthy + name: "{{ volume_name }}" + timeout: 60s + + - action: print + msg: "Volume healthy — bootstrap closure complete" + + # Phase 2: Main workload — fio + dd_write with checksum for data continuity. + - name: record-before + actions: - action: fio_json node: m01 device: "{{ device }}" @@ -129,6 +165,7 @@ phases: bs: 1M count: "2" seek: "16" + sync_mode: fsync save_as: pre_failover_md5 - action: dd_read_md5 @@ -143,10 +180,6 @@ phases: actual: "{{ pre_failover_verify }}" expected: "{{ pre_failover_md5 }}" - - action: wait_volume_healthy - name: "{{ volume_name }}" - timeout: 60s - - action: iscsi_cleanup node: m01 ignore_error: true @@ -156,9 +189,6 @@ phases: - action: print msg: "=== Killing primary ({{ before_server }}) ===" - # The cluster starts m02 before m01, so the natural initial primary is - # the first server (m02 / vs1_pid). Keep the kill target fixed here and - # use discover_primary only as an evidence check. - action: exec node: m02 cmd: "kill -9 {{ vs1_pid }}" @@ -168,14 +198,11 @@ phases: - action: print msg: "Primary killed. Waiting for lease expiry (45s)..." - # Lease TTL is 30s. Wait 45s for expiry + master failover cycle. - action: sleep duration: 45s - name: verify-auto-failover actions: - # Master should auto-promote m01 (the surviving replica) to primary. - # Wait for primary to change from m02 to something else. - action: wait_block_primary name: "{{ volume_name }}" not: "{{ before_server }}" @@ -198,8 +225,6 @@ phases: - name: verify-io-after actions: - # Reconnect iSCSI to the new primary (m01, which is local). - # Use the original lookup vars — iSCSI addr is on the VS, not from registry. - action: iscsi_login_direct node: m01 host: "10.0.0.1" diff --git a/weed/storage/blockvol/testrunner/scripts/run-phase20-t6.ps1 b/weed/storage/blockvol/testrunner/scripts/run-phase20-t6.ps1 index 5edcaea15..aacf0c969 100644 --- a/weed/storage/blockvol/testrunner/scripts/run-phase20-t6.ps1 +++ b/weed/storage/blockvol/testrunner/scripts/run-phase20-t6.ps1 @@ -84,7 +84,7 @@ function Invoke-Pack { } $stage0Pack = @( - (New-Scenario -Id "P20-H0" -Path "weed/storage/blockvol/testrunner/scenarios/internal/recovery-baseline-failover.yaml" -Purpose "Stage 0 bootstrap closure: promoted primary learns replica membership and can reach publish_healthy on the healthy RF=2 sync_all path") + (New-Scenario -Id "P20-H0" -Path "weed/storage/blockvol/testrunner/scenarios/internal/recovery-bootstrap-closure.yaml" -Purpose "Stage 0 bootstrap closure: create -> first fsync fence -> publish_healthy on the healthy RF=2 sync_all path") ) $stage1Pack = @( diff --git a/weed/storage/blockvol/testrunner/suite.go b/weed/storage/blockvol/testrunner/suite.go index eeb43a3da..13fb436b7 100644 --- a/weed/storage/blockvol/testrunner/suite.go +++ b/weed/storage/blockvol/testrunner/suite.go @@ -88,6 +88,7 @@ type SuiteEvidence struct { GlogPatterns []string `yaml:"glog_patterns"` DebugEndpoints []string `yaml:"debug_endpoints"` SaveTo string `yaml:"save_to"` + RunNode string `yaml:"run_node"` // node where scenarios run (for result pull) } // ParseSuiteFile reads and parses a suite YAML file. diff --git a/weed/storage/blockvol/testrunner/suites/phase20-t6-stage0.yaml b/weed/storage/blockvol/testrunner/suites/phase20-t6-stage0.yaml index 3c8656b63..72fb1efc1 100644 --- a/weed/storage/blockvol/testrunner/suites/phase20-t6-stage0.yaml +++ b/weed/storage/blockvol/testrunner/suites/phase20-t6-stage0.yaml @@ -28,7 +28,7 @@ deploy: remote: /opt/work/sw-test-runner scenarios: - - path: weed/storage/blockvol/testrunner/scenarios/internal/recovery-baseline-failover.yaml + - path: weed/storage/blockvol/testrunner/scenarios/internal/recovery-bootstrap-closure.yaml id: P20-H0 evidence: @@ -36,4 +36,5 @@ evidence: debug_endpoints: - "http://10.0.0.1:18480/debug/block/shipper" - "http://10.0.0.3:18480/debug/block/shipper" - save_to: results/phase20-t6/stage0 + save_to: V:/share/sw-block-evidence/phase20-t6 + run_node: m01 diff --git a/weed/storage/blockvol/testrunner/suites/phase20-t6-stage1.yaml b/weed/storage/blockvol/testrunner/suites/phase20-t6-stage1.yaml index 47f84caf1..7a9aacd26 100644 --- a/weed/storage/blockvol/testrunner/suites/phase20-t6-stage1.yaml +++ b/weed/storage/blockvol/testrunner/suites/phase20-t6-stage1.yaml @@ -16,30 +16,27 @@ deploy: goos: linux goarch: amd64 targets: [weed, sw-test-runner] - repo_dir: /c/work/seaweedfs + repo_dir: C:/work/seaweedfs kill_ports: [9433, 18480, 3295] clean_dirs: ["/tmp/sw-fo-*"] binaries: - - local: /c/work/seaweedfs/weed-linux + - local: C:/work/seaweedfs/weed-linux remote: /tmp/sw-test-runner/weed - - local: /c/work/seaweedfs/weed-linux + - local: C:/work/seaweedfs/weed-linux remote: /opt/work/weed - - local: /c/work/seaweedfs/sw-test-runner-linux + - local: C:/work/seaweedfs/sw-test-runner-linux remote: /opt/work/sw-test-runner +# Stage 1: Start with just recovery-baseline-failover (now with 256M WAL). +# Add remaining scenarios after H1A passes. scenarios: - path: weed/storage/blockvol/testrunner/scenarios/internal/recovery-baseline-failover.yaml id: P20-T6-H1A - - path: weed/storage/blockvol/testrunner/scenarios/internal/suite-ha-failover.yaml - id: P20-T6-H1B - - path: weed/storage/blockvol/testrunner/scenarios/cp11b3-manual-promote.yaml - id: P20-T6-H1C - - path: weed/storage/blockvol/testrunner/scenarios/lease-expiry-write-gate.yaml - id: P20-T6-H1D evidence: glog_patterns: ["/tmp/weed.*.INFO.*"] debug_endpoints: - "http://10.0.0.1:18480/debug/block/shipper" - "http://10.0.0.3:18480/debug/block/shipper" - save_to: results/phase20-t6/stage1 + save_to: V:/share/sw-block-evidence/phase20-t6 + run_node: m01 diff --git a/weed/storage/blockvol/wal_shipper.go b/weed/storage/blockvol/wal_shipper.go index 9ebac8c2d..a51e5033b 100644 --- a/weed/storage/blockvol/wal_shipper.go +++ b/weed/storage/blockvol/wal_shipper.go @@ -78,6 +78,10 @@ type WALShipper struct { // Set via SetOnStateChange. Nil = no callback. onStateChange func(from, to ReplicaState) + // onBarrierFailure reports the semantic reason for a failed barrier attempt. + // Used by the host to surface bounded durability failures into diagnostics/core. + onBarrierFailure func(reason string) + // liveShippingPolicy gates whether this shipper may accept current live-tail // WAL entries. The host uses this to keep a replica in bounded catch-up until // the active session contract allows live streaming again. @@ -90,6 +94,11 @@ func (s *WALShipper) SetOnStateChange(fn func(from, to ReplicaState)) { s.onStateChange = fn } +// SetOnBarrierFailure registers a callback for failed barrier attempts. +func (s *WALShipper) SetOnBarrierFailure(fn func(reason string)) { + s.onBarrierFailure = fn +} + // SetReplicaID sets the stable replica identity carried from the host-side // session contract. When empty, transport-level behavior still works but // protocol-aware gating cannot make per-replica decisions. @@ -213,10 +222,16 @@ func (s *WALShipper) CatchUpTo(targetLSN uint64) (uint64, error) { return 0, nil } + log.Printf("wal_shipper: catch-up start replica=%s target_lsn=%d state=%s flushed_lsn=%d data=%s ctrl=%s", + s.replicaID, targetLSN, s.State(), s.replicaFlushedLSN.Load(), s.dataAddr, s.controlAddr) + targetState, replicaFlushedLSN, err := s.reconnectWithHandshake() switch targetState { case ReplicaInSync: s.markInSync() + s.resetCtrlConn() + log.Printf("wal_shipper: catch-up not needed replica=%s handshake_state=%s replica_flushed=%d target_lsn=%d", + s.replicaID, targetState, replicaFlushedLSN, targetLSN) if replicaFlushedLSN > targetLSN { return targetLSN, nil } @@ -224,6 +239,8 @@ func (s *WALShipper) CatchUpTo(targetLSN uint64) (uint64, error) { case ReplicaCatchingUp: achievedLSN, catchErr := s.runCatchUpTo(replicaFlushedLSN, targetLSN) if catchErr != nil { + log.Printf("wal_shipper: catch-up failed replica=%s from_lsn=%d target_lsn=%d err=%v", + s.replicaID, replicaFlushedLSN, targetLSN, catchErr) s.catchupFailures++ if s.catchupFailures >= maxCatchupRetries { s.state.Store(uint32(ReplicaNeedsRebuild)) @@ -233,12 +250,19 @@ func (s *WALShipper) CatchUpTo(targetLSN uint64) (uint64, error) { return achievedLSN, ErrReplicaDegraded } s.markInSync() + s.resetCtrlConn() + log.Printf("wal_shipper: catch-up complete replica=%s achieved_lsn=%d target_lsn=%d", + s.replicaID, achievedLSN, targetLSN) return achievedLSN, nil case ReplicaNeedsRebuild: s.state.Store(uint32(ReplicaNeedsRebuild)) + log.Printf("wal_shipper: catch-up escalated to rebuild replica=%s target_lsn=%d err=%v", + s.replicaID, targetLSN, err) return replicaFlushedLSN, fmt.Errorf("reconnect: %w", err) default: s.markDegraded() + log.Printf("wal_shipper: catch-up left replica degraded replica=%s target_lsn=%d err=%v", + s.replicaID, targetLSN, err) if err != nil { return replicaFlushedLSN, err } @@ -257,21 +281,42 @@ func (s *WALShipper) Barrier(lsnMax uint64) error { } st := s.State() + log.Printf("wal_shipper: barrier start replica=%s state=%s target_lsn=%d flushed_lsn=%d has_progress=%v data=%s ctrl=%s", + s.replicaID, st, lsnMax, s.replicaFlushedLSN.Load(), s.hasFlushedProgress.Load(), s.dataAddr, s.controlAddr) switch st { case ReplicaInSync: // proceed normally to barrier case ReplicaDisconnected, ReplicaDegraded: - if s.hasFlushedProgress.Load() && s.wal != nil { + if s.wal != nil && lsnMax > 0 { + // Integrated bootstrap case: writes may have accumulated before the + // shipper was configured. Replaying the retained prefix up to the + // barrier target closes the "late-configured first fsync" gap. + log.Printf("wal_shipper: barrier recovery via bounded catch-up replica=%s state=%s target_lsn=%d", + s.replicaID, st, lsnMax) + if _, err := s.CatchUpTo(lsnMax); err != nil { + log.Printf("wal_shipper: barrier recovery catch-up failed replica=%s target_lsn=%d err=%v", + s.replicaID, lsnMax, err) + return err + } + } else if s.hasFlushedProgress.Load() && s.wal != nil { // Previously synced — reconnect handshake + catch-up path. + log.Printf("wal_shipper: barrier recovery via reconnect replica=%s state=%s target_lsn=%d", + s.replicaID, st, lsnMax) if err := s.doReconnectAndCatchUp(); err != nil { + log.Printf("wal_shipper: barrier reconnect failed replica=%s target_lsn=%d err=%v", + s.replicaID, lsnMax, err) return err } } else { - // Fresh bootstrap or no WAL access — reset connections for bare retry. + // Fresh bootstrap with no retained target — reset connections for bare retry. + log.Printf("wal_shipper: barrier reset connections for bootstrap retry replica=%s state=%s", + s.replicaID, st) s.resetConnections() } default: // Connecting, CatchingUp, NeedsRebuild — reject immediately + log.Printf("wal_shipper: barrier rejected replica=%s state=%s target_lsn=%d reason=state_not_ready", + s.replicaID, st, lsnMax) return ErrReplicaDegraded } @@ -286,30 +331,22 @@ func (s *WALShipper) Barrier(lsnMax uint64) error { defer s.ctrlMu.Unlock() if err := s.ensureCtrlConn(); err != nil { - s.markDegraded() - s.recordBarrierMetric(barrierStart, true) - return ErrReplicaDegraded + return s.failBarrier("barrier_ctrl_connect_failed", barrierStart, ErrReplicaDegraded) } s.ctrlConn.SetDeadline(time.Now().Add(barrierTimeout)) if err := WriteFrame(s.ctrlConn, MsgBarrierReq, req); err != nil { - s.markDegraded() - s.recordBarrierMetric(barrierStart, true) - return ErrReplicaDegraded + return s.failBarrier("barrier_req_write_failed", barrierStart, ErrReplicaDegraded) } msgType, payload, err := ReadFrame(s.ctrlConn) if err != nil { - s.markDegraded() - s.recordBarrierMetric(barrierStart, true) - return ErrReplicaDegraded + return s.failBarrier("barrier_resp_read_failed", barrierStart, ErrReplicaDegraded) } if msgType != MsgBarrierResp || len(payload) < 1 { - s.markDegraded() - s.recordBarrierMetric(barrierStart, true) - return ErrReplicaDegraded + return s.failBarrier("barrier_bad_response", barrierStart, ErrReplicaDegraded) } resp := DecodeBarrierResponse(payload) @@ -321,8 +358,8 @@ func (s *WALShipper) Barrier(lsnMax uint64) error { // response). This must NOT count as successful sync_all durability because // no authoritative durable progress was established. if resp.FlushedLSN == 0 { - s.recordBarrierMetric(barrierStart, true) - return fmt.Errorf("wal_shipper: barrier OK but no FlushedLSN reported (legacy response)") + return s.failBarrier("barrier_missing_flushed_lsn", barrierStart, + fmt.Errorf("wal_shipper: barrier OK but no FlushedLSN reported (legacy response)")) } // Barrier success with durable progress — transition to InSync. s.markInSync() @@ -338,23 +375,21 @@ func (s *WALShipper) Barrier(lsnMax uint64) error { } } s.recordBarrierMetric(barrierStart, false) + log.Printf("wal_shipper: barrier success replica=%s target_lsn=%d flushed_lsn=%d", + s.replicaID, lsnMax, resp.FlushedLSN) return nil case BarrierEpochMismatch: - s.markDegraded() - s.recordBarrierMetric(barrierStart, true) - return fmt.Errorf("wal_shipper: barrier epoch mismatch") + return s.failBarrier("barrier_epoch_mismatch", barrierStart, + fmt.Errorf("wal_shipper: barrier epoch mismatch")) case BarrierTimeout: - s.markDegraded() - s.recordBarrierMetric(barrierStart, true) - return fmt.Errorf("wal_shipper: barrier timeout on replica") + return s.failBarrier("barrier_timeout", barrierStart, + fmt.Errorf("wal_shipper: barrier timeout on replica")) case BarrierFsyncFailed: - s.markDegraded() - s.recordBarrierMetric(barrierStart, true) - return fmt.Errorf("wal_shipper: barrier fsync failed on replica") + return s.failBarrier("barrier_fsync_failed", barrierStart, + fmt.Errorf("wal_shipper: barrier fsync failed on replica")) default: - s.markDegraded() - s.recordBarrierMetric(barrierStart, true) - return fmt.Errorf("wal_shipper: unknown barrier status %d", payload[0]) + return s.failBarrier("barrier_unknown_status", barrierStart, + fmt.Errorf("wal_shipper: unknown barrier status %d", payload[0])) } } @@ -364,6 +399,21 @@ func (s *WALShipper) recordBarrierMetric(start time.Time, failed bool) { } } +func (s *WALShipper) notifyBarrierFailure(reason string) { + if s.onBarrierFailure != nil { + s.onBarrierFailure(reason) + } +} + +func (s *WALShipper) failBarrier(reason string, start time.Time, err error) error { + s.markDegraded() + s.recordBarrierMetric(start, true) + s.notifyBarrierFailure(reason) + log.Printf("wal_shipper: barrier failed replica=%s reason=%s target_flushed=%d err=%v data=%s ctrl=%s", + s.replicaID, reason, s.replicaFlushedLSN.Load(), err, s.dataAddr, s.controlAddr) + return err +} + // ShippedLSN returns the highest LSN successfully sent to the replica (diagnostic only). // This is NOT authoritative for sync durability — use ReplicaFlushedLSN() instead. func (s *WALShipper) ShippedLSN() uint64 { @@ -496,18 +546,33 @@ func (s *WALShipper) resetConnections() { s.ctrlMu.Unlock() } +func (s *WALShipper) resetCtrlConn() { + s.ctrlMu.Lock() + if s.ctrlConn != nil { + s.ctrlConn.Close() + s.ctrlConn = nil + } + s.ctrlMu.Unlock() +} + // doReconnectAndCatchUp runs the full reconnect handshake + catch-up protocol. // On success, transitions to InSync and resets ctrl connection for barrier. func (s *WALShipper) doReconnectAndCatchUp() error { + log.Printf("wal_shipper: reconnect start replica=%s state=%s flushed_lsn=%d data=%s ctrl=%s", + s.replicaID, s.State(), s.replicaFlushedLSN.Load(), s.dataAddr, s.controlAddr) targetState, replicaFlushed, err := s.reconnectWithHandshake() switch targetState { case ReplicaInSync: s.markInSync() + log.Printf("wal_shipper: reconnect complete replica=%s state=%s replica_flushed=%d", + s.replicaID, targetState, replicaFlushed) case ReplicaCatchingUp: // Use the handshake-reported flushedLSN as catch-up start point, // NOT the shipper's cached value. The replica may have lost progress // since the shipper last heard from it. if catchErr := s.runCatchUp(replicaFlushed); catchErr != nil { + log.Printf("wal_shipper: reconnect catch-up failed replica=%s from_lsn=%d err=%v", + s.replicaID, replicaFlushed, catchErr) s.catchupFailures++ if s.catchupFailures >= maxCatchupRetries { s.state.Store(uint32(ReplicaNeedsRebuild)) @@ -517,20 +582,21 @@ func (s *WALShipper) doReconnectAndCatchUp() error { return ErrReplicaDegraded } s.markInSync() + log.Printf("wal_shipper: reconnect catch-up complete replica=%s from_lsn=%d", + s.replicaID, replicaFlushed) case ReplicaNeedsRebuild: s.state.Store(uint32(ReplicaNeedsRebuild)) + log.Printf("wal_shipper: reconnect escalated to needs_rebuild replica=%s err=%v", + s.replicaID, err) return fmt.Errorf("reconnect: %w", err) default: s.markDegraded() + log.Printf("wal_shipper: reconnect left replica degraded replica=%s err=%v", + s.replicaID, err) return ErrReplicaDegraded } // Reset ctrl connection so barrier creates a fresh one. - s.ctrlMu.Lock() - if s.ctrlConn != nil { - s.ctrlConn.Close() - s.ctrlConn = nil - } - s.ctrlMu.Unlock() + s.resetCtrlConn() return nil }