mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-06 06:22:05 +02:00
f4ef37e7526fcfdb9ae16ab83bed02433ffbd804
245
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
10b0f2b8ad |
volume server: refuse the rest of a grouped run after a durable index failure (#11576)
* volume server: refuse the rest of a grouped run after a durable index failure A durable write whose needle-map put fails stops the volume taking writes (#10825): sent on its own, the next write then fails read only before it appends. The grouped run from #11543 appends and syncs every entry before publishing any, then kept publishing the entries after the failed one and acked them once the shared .idx sync went through. When the failed put tore its .idx row, the rows appended after it land off alignment, so the next load parses them as garbage and the acked writes are gone. Once a durable entry fails to publish, refuse every later entry of the run with ReadOnly, as the per-needle path does. The entries before it stay acked; their rows go down with the run's one .idx sync. The refused records stay on the .dat unindexed, as the failed one does on its own. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: refuse a grouped entry staged as a cookie mismatch too After a durable entry in a grouped run fails to index, the entries after it are refused as they would be on their own. On its own an entry meets check_writable before its cookie check, so one staged as a cookie mismatch now gets the refusal too, instead of keeping its staging error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: trim a torn .idx row back so the next stays aligned A failed write_index_entry can leave half a row in the .idx. With the writer appending at the tail, every row written after it lands off alignment and the next load parses them as garbage, so a write acked behind a torn row does not come back. Trim the file back to idx_file_offset on a failed append, in both needle maps, and cover it with a test that writes past a torn row and reloads. * volume server: refuse queued Go writes once a durable index update fails processBatch kept writing after a failed nm.Put, and the single-write path checked IsReadOnly only outside the volume lock. A durable write whose index update fails now marks the volume noWriteOrDelete, and each queued request is checked before it appends, so the ones after a failed durable entry are refused the way a lone write is. Deletes get the same noWriteOrDelete refusal a lone delete gets. * volume server: refuse appends while a torn .idx row cannot be trimmed When trimming back a half-written .idx row itself fails, the next append would land after the torn bytes and every later row would parse off alignment on load. Latch the map as torn and refuse appends until the trim succeeds, on both CompactNeedleMap and RedbNeedleMap; the same latch covers an orphan row that could not be trimmed after a failed redb commit. The .idx writer is now opened with write+append access so truncate_to (set_len) works on Windows, where an append-only handle cannot trim. * volume server: write .idx rows at idx_file_offset, not via append mode Rust's OpenOptions on Windows strips FILE_WRITE_DATA whenever append is set so the handle stays strictly append-only, which makes set_len fail - the torn-row trim could never succeed there. Open the .idx writer with plain write access and seek to idx_file_offset before each row, the same positioned-write model the Go server uses. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
8c1ebbee32 |
volume server: split get_or_head_handler_inner into phases (#11489)
* volume server: read GET/HEAD needles off the store lock, and only once The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stream needle chunks without the store lock StreamingBody::poll_frame took store.read() and find_volume for every chunk to compare the volume's compaction revision, dup'd the source handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the stream also holds a data-file read lease for its whole life, while a writer waits for that lease under store.write(): the next chunk's store.read() then waits for the writer and the writer for the stream. The per-chunk re-lookup was also wrong. The stream reads a handle opened at plan time, which pins the .dat inode the offset was resolved against; a vacuum commit renames a new file over .dat and leaves that inode untouched. The re-looked-up offset belongs to the new file but was read from the old inode, so a stream whose needle a vacuum moved ended in a checksum error. The pinned offset stays valid, so the check, and with it every store access, is dropped, along with the now unused re_lookup_needle_data_offset and the revision fields of the read plan. The source is shared as an Arc instead of dup'd per chunk, and the chunk buffer is a BytesMut that the blocking read hands back with its result, so its allocation is reclaimed once the previous frame has been written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: split get_or_head_handler_inner into phases get_or_head_handler_inner was a ~650-line function. Its middle resolved the needle and set five mutable flags (stream_info, can_stream, can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that three if-let reply paths then re-tested, each re-checking stream_info. It is now a 126-line orchestrator over named phases: reject_read_jwt, proxy_missing_volume, wait_for_download_slot, parse_read_request, read_ec_needle / read_volume_needle, etag_and_last_modified, not_modified_response, read_response_headers, and the reply phases stream_response, head_from_meta_response, range_from_source_response, buffered_payload and buffered_response. The read phases return a ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource, Buffered) carries the NeedleStreamInfo only on the variants that use it, so the reply is one match instead of three flag checks. Pure refactor: every status code, header and header order, error text, metric increment, lock and data-file lease scope, spawn_blocking boundary and side-effect order is unchanged. Phases that can end the request return ControlFlow<Response, T>. A Range header that is not visible ASCII still falls through to the buffered path, as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a needle stream once its volume becomes unavailable Taking the store lock out of StreamingBody also dropped its per-chunk unavailable_error() check. With -hasSlowRead a writer can take the data-file lease between chunks, fail its fsync and its truncate, and mark the volume unavailable; the stream then kept serving the rest of the needle from its pinned handle. The volume's io_unavailable reason is now an Arc-shared leaf mutex that the read plan hands to the stream. Each chunk checks it under its data-file lease, where the writer marks it, and fails with the same "volume is unavailable: <reason>" error the old check returned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: mirror Go order in the buffered read path - check HEAD before Range in buffered_response (writeResponseContent order); an EC-volume HEAD with a Range header answered 206, Go answers 200 - treat the proxied flag as an exact query pair like Go's parsed lookup, not a substring - name the phases after their Go counterparts: check_download_limit and read_ec_shard_needle; reuse has_replication() - drop comments that restate the code or cite Go line numbers --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
07da302da0 |
volume server: ec.decode verifies, cleans up and compacts like Go, off the runtime (#11547)
* volume server: ec.decode reads the .ecx from the index dir it was copied to VolumeEcShardsCopy writes the .ecx/.ecj into the receiver's -dir.idx, so with a split data/index dir the decode target has no .ecx beside its shards. VolumeEcShardsToVolume sized the .dat from the right .ecx but built the .idx from the data dir, failing with NotFound after the .dat was already published. It now reads .ecx/.ecj from where the EC volume opened them and writes the .idx beside the .dat, where Go leaves it. The live-entry check and the .dat size also ignored deletions recorded only in the .ecj, which Go folds into the .ecx (RebuildEcxFile) first: a fully deleted volume was decoded instead of reported as having no live entries, and deleted tail needles were copied into the .dat. Both now treat journaled ids as deleted, without rewriting the sealed .ecx. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode keeps the decoded volume writable and reads every .ecj The rebuilt .idx copied a journaled tail needle's .ecx row verbatim after the .dat was cut short before it, so the mount saw a row past EOF and marked the decoded volume read-only. Rows of deleted needles the .dat no longer holds are now dropped, and each journaled needle still in the .dat gets one tombstone instead of one per journal entry. VolumeEcShardsCopy appends journals collected from other holders into the idx dir, but the decode read only the .ecj beside the .ecx, which sits in the data dir when this server generated the shards. It now reads both, once, in bounded chunks via the loader EcVolume uses. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: test ec.decode drops a sealed .ecx tail tombstone Covers the other half of the rule added in the previous commit: a tail needle tombstoned in the .ecx itself (Go's RebuildEcxFile) is cut from the .dat, and its row must not reach the rebuilt .idx either. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode runs its file I/O off the async runtime VolumeEcShardsToVolume released the store lock before decoding, but read the .ecx/.ecj, rebuilt the .dat and wrote the .idx inside the async handler, parking a runtime worker for the length of a volume-sized copy. The decode now runs in spawn_blocking on inputs snapshotted under the store lock. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode checks the rebuilt .dat is complete Go stats the decoded .dat before writing the .idx (VerifyDecodedDatFile) and fails the decode when it is shorter than the extent the EC index references, since the caller deletes the shards once the call returns. The Rust handler returned success without that check. The rebuild already fails on a short shard read, so this guards the published file itself. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode drops the decoded volume's bitrot sidecars Go removes <base>.ecsum and <base>.ecsum.v<N> beside the .dat and beside the .ecx once the .idx is written, so a stale checksum sidecar cannot pass for the protection of a later re-encode. The Rust handler left them in place. Removal is best effort, as in Go. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode compacts the decoded volume Go ends VolumeEcShardsToVolume with an offline CompactVolumeFiles, so the decoded volume holds only live needles. The Rust decode left every needle deleted through the .ecj in the .dat, tombstoned in the .idx, until a later vacuum reclaimed it. Store::compact_volume_files loads the unmounted volume, checks free space the way the vacuum does (the estimate now lives in one helper), and runs the vacuum's compact-by-index and commit. As in Go a failed compaction is logged and the decode still succeeds, so the uncompacted .idx rules stay: the tests that pin them now make the compaction fail. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode keeps deletes journaled while the .dat is written The decode read the .ecj journals once, before rebuilding the .dat, so a delete that reached the EC volume during the rebuild was left out of the new .idx and the needle came back live. Each journal's read length is now kept, and the bytes appended since are read just before the .idx is written, after waiting out any journal append in flight (appends hold the store write lock), so every delete acknowledged by then is in the .idx. A delete after that point is still lost, as in Go. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Guard overlapping ec decode requests; serialize journal catch-up volume_ec_shards_to_volume runs its decode in spawn_blocking, so a dropped request leaves the job running and a retry would race it on the temporary and final volume files. Claim the vid in a per-server in-flight set until the blocking job finishes, and return Unavailable to an overlapping request. The Go handler has the same exposure and gets the same guard. Journal appends hold the store write lock through their sync-or-truncate, so holding a read lock across the catch-up read guarantees every record it sees is committed: a rolled-back delete can no longer leave a tombstone in the decoded index. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Reconcile the swap when offline compaction commit fails A CommitCompact that fails after the .cpc marker may have renamed .dat but not .idx. cleanup_compact refuses while the marker exists, so the mismatched pair survived until a restart reconciled it — and the decode caller treats the failure as non-fatal. Run reconcileCompactState on commit failure so a decided swap rolls forward and orphan temps are removed before the volume can mount. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Release the decode claim on panic * volume: add ec_decodes_in_flight to the integration-test state literal * volume server: hold the decode tail's lock through compaction The catch_up read released before the rebuilt .idx was written and the volume compacted, so a delete synced to .ecj in that window was durably journaled yet absent from the published index — resurrecting the needle. Rust now holds the store read lock from catch_up through compact, and Go mirrors it by holding the volume's journal lock from the journal- consuming index write through CompactVolumeFiles. * volume server: serialize ec decode's tail per volume, not per store Review follow-ups on the decode path: - Rust: holding the store read lock from journal catch-up through the offline compaction stalled every writer on unrelated volumes for the whole rewrite. The new ec_decode_tail set marks the vid only while its .idx is published and .cpd/.cpx swapped; the two local .ecj append paths (VolumeEcBlobDelete, the distributed delete's local journal) wait on a Notify for that span — Go's per-volume ecjFileAccessLock semantics without the global stall. VolumeMount and the staged-adopt path are also held off while a decode claim is in flight so neither can race the swap. - Rust: the initial journal read ran unlocked, so bytes a rolled-back append later truncated could be folded in as phantom tombstones. The first pass stays unlocked (a slow journal must not stall the store) and a rescan under the quiescing read lock re-reads only committed content; catch_up now rebuilds the id set when a regular journal shrank. - Go: the decode resolved the compaction DiskLocation through FindEcVolume while holding the journal lock, inverting DestroyEcVolume's map->journal order into a deadlock. The lookup now happens first, and DestroyEcVolume/deleteEcVolumeById/DiskLocation.Close destroy outside the map lock. - Go: RebuildEcxFile unlinks .ecj while the volume's ecjFile handle stays open, so later deletes could commit to a detached inode. Both call sites now fold under the journal lock and ReopenDeletionJournal repoints the handle at the live path, working on the volume's resolved .ecx dir (EcIndexBaseFileName) rather than the configured index dir. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: fence EC remounts behind the destroy tombstone DestroyEcVolume, deleteEcVolumeById, and the collection-delete sweep now remove the EcVolume from ecVolumes before destroying it off-lock, so a concurrent remount could re-open shard files that the in-flight destroy then unlinks — registering a detached fd. Each destroy records a per-vid tombstone channel in a new ecVolumesDestroying map before dropping the map entry and closes it when Destroy returns. The tombstone intentionally survives as the vid's destroy generation: loadEcShardWithIdxDir compares it before and after opening the shard, so a destroy that both started and finished inside the open window is still detected. A mismatch drops the just-opened shard (releasing its fd and mount gauge) and retries after the destroy completes; a successful mount clears the stale tombstone. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: rescan the .ecj under the store lock only after a rollback The decode's second journal pass ran a full rescan under the store read lock on every decode, stalling unrelated writers for the length of the scan. Bump a process-wide epoch whenever a failed append truncates its uncommitted tail; an unchanged epoch between the unlocked read and the quiesced pass proves every id folded in was committed, so catch_up() suffices. catch_up() also treats a journal that was read but has since disappeared as shrunk to zero, so its earlier ids cannot linger. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: check the decode tail under the store write lock on delete A blob delete waited for the publishing tail before taking the store write lock, so a decode that claimed the tail while the delete was parked behind the decoder's read lock could still see the journal append land after the rebuilt .idx — an acknowledged delete the mount would miss. Test tail membership under the write lock instead, retrying after the wait; journal_delete_local reports WouldBlock for the same recheck on the distributed path. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: claim the vid for mount and staged adoption, per volume VolumeMount and the staged .copying adoption held the ec_decodes_in_flight set lock through slow file renames and mounts, stalling every unrelated volume's decode, mount, and adoption. Take the per-volume claim instead — the same exclusion against a racing decode for this vid, released when the call returns. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: fail the decode when a compaction commit marker survives CompactVolumeFiles' caller logged a compaction error and went on to delete the EC shards. When the commit marker (.cpc) is still on disk the .dat/.idx swap was decided but could not be reconciled, so the mounted pair may be mismatched — report the failure instead so the shards are kept and the caller can retry. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: gate the parked-delete test on the held write lock The releaser thread and the spawned delete raced for the store write lock; on a slow runner the delete could acquire it first and commit before the tail was ever claimed, failing !delete.is_finished() on the Windows unit-test job. Spawn the delete only after the thread reports the lock held. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
0ff7794c54 |
volume: compact an oversized .ecj at mount, safely (Rust + Go) (#11555)
* volume: compact an oversized .ecj at mount, safely (Rust + Go) Restore the mount-time compaction dropped from #11408, Rust + Go parity. A journal already bloated by repeated shard copies is folded down to the id set it encodes. - Trigger after load when file_records > max(threshold, 4x distinct), with a 1 MiB floor so small journals are never rewritten. The set is written to .ecj.compact.tmp + fsync, the handle dropped, renamed, the directory fsynced and the append handle reopened. A failure before the rename keeps the original journal and handle; a failure after it fails the mount. - Go never compacts after a failed journal load; the set would be partial and the rewrite would drop the unread records. - A per-path registry (ecj_registry.rs / ecj_registry.go) counts EcVolume holders and out-of-band writers of each .ecj. Compaction runs only when this volume is the sole holder and no copy is writing; holders and writers wait while one runs. This covers shared -dir.idx journals and cross-disk reconcile, where another EcVolume may hold the same journal. - VolumeEcShardsCopy and EC index recovery register as writers around their .ecj append and partial-file cleanup. - Under the reservation, re-check that the file on disk is still the inode and size that was loaded. - Publish errors are classified where they happen; a failed rename plus a failed restore reports both errors. - Compaction runs after the .vif / bitrot checks, so a refused mount leaves the journal untouched. - The tmp is opened like other volume files, removed at mount if a crash left it, and listed in every EC index cleanup path. Failure paths are tested through the real mount via injectable fs steps (open_with / newEcVolumeWith), plus sibling holders, active copies, changed-after-load, stale tmp cleanup, refused mounts and the Go load-error guard. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: fail the mount when the compacted .ecj's directory cannot be synced The Rust mount synced the journal's directory after renaming the compacted file over it through the crate's best-effort fsync_dir, which returns Ok when the directory cannot be opened. A rename needs only write and search permission, so on a directory without read permission the replacement was published, never synced, and the mount went on taking deletes against it. Sync through a helper that propagates the open error, as Go's util.FsyncDir already does, so that case fails the mount like any other post-rename sync failure. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: test the no-compaction-after-failed-load rule through the Go mount The test for it handed compactEcjAfterLoad an artificial error on a volume that had loaded cleanly, so it would not notice NewEcVolume dropping the real load error on the way to compaction. Make the journal read one of the injectable ecjFsOps steps and fail it inside the real mount, after the first chunk, on a journal whose last entry is an id the first chunk does not hold. The mount must leave the file byte for byte as it was; a clean remount then compacts and keeps that id. The Rust mount fails outright on a load error, so it has no equivalent path. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: register ReceiveFile's .ecj writes with the journal registry ReceiveFile refuses a mounted EC volume only once, when the info message arrives, then creates the .ecj and streams chunks into it. A volume that mounted on that journal mid-stream could find a bloated prefix, pass the inode-and-size re-check and rename a compacted file over it; the rest of the stream then went to the unlinked inode and was lost. Register the path as a writer before the file is created, in both the Go and Rust handlers, and hold it until the file is closed and any partial copy removed, as the shard-copy and index-recovery appends already do. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: skip .ecj compaction when a writer ran since the journal was loaded Compaction checked only that no writer was active at the reservation, and that the file was still the loaded inode at the loaded size. A ReceiveFile truncates and refills the journal in place, so one that ran during the mount's load, or after it, and finished before the reservation could leave different ids at the same length; compaction then wrote the stale set over them. Give each path a write generation that every writer bumps as it starts. A holder records it, and whether a writer was active, when it registers, which is before it opens and loads the journal. It may compact only if no writer was active then and the generation has not moved. Same rule in Go and Rust; the journal read becomes an injectable step in Rust as it is in Go, so both test the in-place rewrite through the real mount. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: match the ReadOnly(VolumeId) variant in write_volume_needles #11543 matched VolumeError::ReadOnly as a unit variant in Store::write_volume_needles, and #11544 changed it to ReadOnly(VolumeId) in the same merge window. Each passed CI on its own, but master no longer compiles the Rust volume server. Carry the volume id through. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
fa77cde7da |
vacuum: check compaction space against live bytes, not volume size (#11524)
* vacuum: size the compaction space check by live bytes, not volume size ensureCompactVolumeSpace required the volume's current .dat and .idx size as free space before compacting. That is the size of the garbage, not of what compaction writes, so on a disk that filled up until its volumes went read-only every compaction was refused, including all-garbage volumes that would compact to a superblock and an empty index. The sweep then retried every volume each cycle and reclaimed nothing (issue #11516). Estimate the output from what the needle map already tracks: live content bytes plus a per-needle framing upper bound behind a superblock, and one index entry per live needle. The estimate never exceeds the current volume size and preallocate still wins when larger. Volumes whose deleted sizes are unknown (.sdx converted back to .idx) keep the whole volume as the estimate. The disk probe moves behind a package variable so the tests can stand in for a full disk; the tests build real volumes instead of re-implementing the formula. * vacuum: space check reserves the index on top of preallocate, checks a separate index disk Review follow-ups: preallocate only stands in for the new .dat, so the rebuilt index is added on top of it; with separate index directories the data disk is checked for the .cpd and the index disk for the .cpx; and the estimates carry 1/16 headroom because counters rebuilt from an index file pass through a Bloom filter with a 0.1% false positive rate. Neither estimate exceeds the current file. * vacuum: split the space check by filesystem, not by directory name Two directories can sit on one filesystem and share its free space, so the data and index estimates are checked separately only when the index directory is on another device; otherwise the sum must fit. Unknown is treated as shared. * vacuum: ask the index directory for its share even when it looks like the same filesystem A volume mounted under the data directory's drive letter on Windows has the same volume name, so the identity check calls it shared. Checking the index directory for the index estimate as well costs one statfs and catches a full index mount either way. * vacuum: identify a Windows volume by its GUID, not its path prefix A volume can be reached through a drive letter and through a folder it is mounted on, so filepath.VolumeName says nothing about the free-space pool. Resolve each directory to its mount point and compare the volume GUIDs; when that fails the two are treated as shared. * vacuum: keep the framing and disk_space_low coverage the rebase displaced * rust volume: split the compaction space check across data and index disks Mirror the Go check: estimate the new .dat and rebuilt .idx separately — live content plus per-needle framing capped at the current file, with preallocate standing in for the data file when larger — and check each directory against its own filesystem's free space. Two directories on one filesystem are asked for the sum. * vacuum: tighten comments on the compaction space check Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
35b090a4df |
volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go) (#11554)
* volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go) An EC volume's deletion journal is a set of needle ids, but shard copy and index recovery appended the peer's whole journal, doubling the file on every ec_balance round trip. Fold the peer's ids in as a union instead: only ids the local journal lacks are appended. - The journal is never replaced. A mounted EcVolume merges a peer's ids through its live handle under the lock deletes take (Go MergeJournal / Rust merge_journal), wherever its journal lives. - An unmounted journal gets only the missing ids appended while mounts are excluded; the delta is read outside the lock and re-read if the journal changed. - The source .ecj streams into memory as an id set: no staging files, chunked reads, memory proportional to distinct ids. - Go and Rust agree that a source journal exists when it sends a modified time or any bytes. A missing source stays a no-op. - Rust runs every merge in spawn_blocking and shares one receive/merge path between shard copy and index recovery. The decode path and the journal format are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: route .ecj merges to the runtime that holds the journal open Disks sharing one index directory all resolved as the journal's owner, so the last one won and a sibling's mounted runtime was skipped: the merge appended behind its open handle and the sibling kept serving the peer's deleted needles until remount. Callers now name the receiving disk by its data directory; the merge goes through that disk's runtime, else a sibling runtime whose journal is the target file. In Go the unmounted append now holds every disk's EC lock (in location order) while it rechecks for a mount, so a sibling mounting from this disk's index during the unlocked read is merged through instead. In Rust a mount that lands during the read is merged through directly and its added count returned, rather than discarded and reported as zero. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: sync merged .ecj records outside the disks' EC locks The unmounted merge held every disk's EC read lock across its fsync, so a slow sync on one disk held off mounts on all of them, along with the EC reads queued behind those mounts. Mounts only need to be excluded while the records are written: the write now happens under the locks and the fsync after they are released, since a later mount reads the written records from the page cache. A failed fsync rolls back only if nothing has mounted the journal or appended to it since the write. A merge through a mounted volume now keeps only that volume's disk locked across its fsync. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: roll back an unsynced .ecj merge through a volume mounted mid-sync If a volume mounted after the unmounted merge wrote its records but before the fsync failed, the rollback kept the records because the journal was now open, leaving ids in the volume's deleted set that may never reach disk; a retried merge then saw them and synced nothing. The rollback now goes through that volume the way its own failed journal fsync does: truncate back and drop the ids from the in-memory set, so a retry appends and syncs them again. It still keeps the records if the volume journaled since, as truncating would lose that delete. No fsync runs under the disk locks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: decide .ecj merge rollback from the journal's actual length Two runtimes can hold one journal (cross-disk mounts). The rollback of an unsynced merge checked one runtime's cached ecjFileSize, which another runtime's appends leave stale, so it could truncate a delete that runtime had already synced. The rollback now holds every holder's journal lock and truncates only if the file's actual length is still the append's end, then updates each holder's size and deleted set. Otherwise later records follow the merged ones, so they stay and are rewritten in place and synced outside the locks, rather than left possibly not durable. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: keep unsynced .ecj merge ids out of mounted deleted sets When a merge's fsync failed, later records blocked the rollback, and the rewrite-and-sync failed as well, the merged ids stayed in every mounted volume's deleted set without being shown durable, so a retried merge saw them as present and synced nothing. They now leave those sets while the records stay in the file, matching DeleteNeedleFromEcx, which publishes an id only after its record syncs. The merge returns the error and a retry appends and syncs them again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: publish merged .ecj ids to every holder of the journal Two runtimes can journal into the same file when disks share an index directory. The merge went through only the first holder, leaving a sibling's in-memory deleted set without the ids, so it could keep serving a needle the peer deleted until it remounted. Every holder of the journal now gets the merged ids, in Go and in the volume server. Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Publish merged .ecj ids to the journal actually written mountedEcJournal prefers the receiving disk's own runtime for the vid, whose journal may live in its data directory while the copied records name a sibling's journal in the index directory. Publishing by the requested ecjPath then marked a holder of a different file deleted on records that file never persisted, resurrecting the needles on remount. Publish by the picked runtime's journal path instead. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
d8f926cf46 |
volume server: stream needle chunks without the store lock (#11488)
* volume server: read GET/HEAD needles off the store lock, and only once The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stream needle chunks without the store lock StreamingBody::poll_frame took store.read() and find_volume for every chunk to compare the volume's compaction revision, dup'd the source handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the stream also holds a data-file read lease for its whole life, while a writer waits for that lease under store.write(): the next chunk's store.read() then waits for the writer and the writer for the stream. The per-chunk re-lookup was also wrong. The stream reads a handle opened at plan time, which pins the .dat inode the offset was resolved against; a vacuum commit renames a new file over .dat and leaves that inode untouched. The re-looked-up offset belongs to the new file but was read from the old inode, so a stream whose needle a vacuum moved ended in a checksum error. The pinned offset stays valid, so the check, and with it every store access, is dropped, along with the now unused re_lookup_needle_data_offset and the revision fields of the read plan. The source is shared as an Arc instead of dup'd per chunk, and the chunk buffer is a BytesMut that the blocking read hands back with its result, so its allocation is reclaimed once the previous frame has been written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a needle stream once its volume becomes unavailable Taking the store lock out of StreamingBody also dropped its per-chunk unavailable_error() check. With -hasSlowRead a writer can take the data-file lease between chunks, fail its fsync and its truncate, and mark the volume unavailable; the stream then kept serving the rest of the needle from its pinned handle. The volume's io_unavailable reason is now an Arc-shared leaf mutex that the read plan hands to the stream. Each chunk checks it under its data-file lease, where the writer marks it, and fails with the same "volume is unavailable: <reason>" error the old check returned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
8a9563e53d |
volume: test that only the disk holding the replaced replica gets the free-slot credit (#11486)
* volume server: split volume_copy into phases and type the delete-after-status gate volume_copy was one ~400-line handler, and the rule that an existing local replica is deleted only after the source's ReadVolumeFileStatus succeeded was held by statement order alone. The keep_remote_data=true that the pre-copy delete and the failed-copy rollback must share was kept in sync by a comment pointing from one to the other. The handler is now a ~60-line orchestrator over connect_to_copy_source, SourceVolumeStatus::fetch, delete_existing_replica, plan_copy_destination and a VolumeCopyJob whose run() drives preallocate_dat, transfer_files, finish_copied_files and mount_and_reply, with cleanup_failed_copy on error. delete_existing_replica takes a &SourceVolumeStatus, which only fetch can construct (private field in a child module), so the delete cannot be called before the status RPC. Both deletes go through delete_replica_keep_remote. Pure refactor: call order, status codes and messages, cancellation checks, throttling, progress reports and cleanup are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: find space for a VolumeCopy before deleting the replica it replaces VolumeCopy deleted an existing local replica as soon as the source answered ReadVolumeFileStatus and only then looked for a location with room for the copy. With no usable location (disk full, low-disk, wrong disk type) the call errored after the delete, leaving the node with neither the old replica nor the new one. Plan the destination first, as Go does: find_free_location_replacing credits the location holding the replaced volume with that volume's slot, so a disk at its volume limit that holds the replica still accepts the copy. Only then delete the replica and write the .note (still after the delete, as in Go). delete_existing_replica now takes the planned CopyDestination, so the delete cannot precede the plan. find_free_location_predicate keeps its behaviour. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: describe the replace-credit test against the current VolumeCopy flow Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
eeec9ec09a |
volume server: refuse to compact a volume tiered to remote storage (#11545)
* volume server: refuse a tier move while compacting, and a commit once tiered A tier move to remote and a vacuum compaction of the same volume could interleave and leave the volume unreadable: - A compaction committing while the upload ran swapped .dat/.idx under the transfer, which reopens the .dat by path per part. The move then published an object holding the old (or a mixed) layout against the compacted .idx, and with keep_local_dat_file=false deleted the only compacted .dat. - A tier move finishing while the compaction copy ran (or between the copy and the commit) let the commit swap in the compacted .idx while the reload served the pre-compaction remote object through it. The tier move now refuses to start while the volume is compacting, and re-checks the compaction revision under the store write lock before it records the remote file; on a mismatch it deletes the uploaded object and fails with FailedPrecondition, leaving the volume local. Committing a compaction on a volume that has a remote file is refused and its .cpd/.cpx removed, since the reload would read the remote object through the compacted index. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: abort a tier move whose volume was replaced or removed The tier-up bookkeeping looked the volume up by id only and compared the compaction revision. A delete and re-create of the same id during the upload yields a fresh volume at the same revision, so the move recorded the old volume's object on the new one and, without keep_local_dat_file, removed the new .dat. An unmounted volume was skipped and the move reported success, leaving the uploaded object referenced by nothing. Capture the volume instance (its data-file access control Arc, as the scan and read plans do) with the revision, and require both under the store write lock. A replaced volume fails with FailedPrecondition, a missing one with NotFound; either way nothing is recorded and the object is deleted after the lock is released. Go fails in both cases because deleting or unmounting closes the descriptor its copy reads. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: refuse to compact a volume tiered to remote storage Committing a compaction of a tiered volume is refused, since the reload would read the remote object through the compacted index. The compaction itself still started: a tiered volume's data backend is the remote object (the local .dat is dropped or deleted on tier-up), so an explicit vacuum streamed the whole .dat out of remote storage into a .cpd that the commit then discarded. Refuse at the start of the compaction instead, before the .cpd is created, at the point where Go's copy opens the local .dat. The truncated-index test now uses a read-only local volume for its sorted index, since a tiered one no longer reaches the copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: match the ReadOnly(VolumeId) variant in write_volume_needles #11543 matched VolumeError::ReadOnly as a unit variant in Store::write_volume_needles, and #11544 changed it to ReadOnly(VolumeId) in the same merge window. Each passed CI on its own, but master no longer compiles the Rust volume server. Carry the volume id through. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
6f9becaa37 |
volume server: answer BatchDelete on EC needles as Go does (#11541)
* volume server: VolumeNeedleStatus reads remote EC shards and reports deleted needles like Go
For an EC volume the handler read only locally mounted shards, so a node
that did not hold the shard with the needle's bytes answered Internal
"ec shard N not available locally". Go's ReadEcShardNeedle fetches the
interval from a peer or reconstructs it. It also mapped every regular
volume read error, including a tombstone, to NotFound "needle not found",
which fs.verify treats as a missing needle; Go returns ErrorDeleted as a
plain error ("already deleted"), which fs.verify skips.
The EC branch now drops the store guard and uses the distributed EC read
the HTTP GET path uses. Errors map like Go: needle absent -> NotFound
"needle not found <decimal id>", tombstoned (regular or EC .ecx/.ecj) ->
Unknown "already deleted", anything else -> Unknown with the error text.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* volume server: tell EC deletions and vanished volumes apart in VolumeNeedleStatus
The distributed EC reader returned Ok(None) for an absent needle, a needle
a peer reported deleted, and a volume unmounted after the handler's own
existence check. VolumeNeedleStatus answered all three NotFound "needle not
found", which fs.verify -pruneEntries counts as lost data. A reported
deletion was also lost when an earlier interval failed.
The reader now says why it has no needle (EcMiss: NotFound, Deleted,
VolumeNotFound), classifying the local tombstone itself and letting a
reported deletion outrank other interval errors, as Go's ReadEcShardNeedle
does. VolumeNeedleStatus maps Deleted to Unknown "already deleted" and
VolumeNotFound to "volume not found", and drops its separate EC pre-check.
read_ec_shard_needle_distributed keeps its Ok(None) for every miss, so the
other callers are unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* volume server: answer BatchDelete on EC needles as Go does
With skip_cookie_check, which every weed/ client sends, an EC needle that
was already deleted came back 404 "ec needle <fid> not found". Go's
DeleteEcShardNeedle gets ErrorDeleted from its read and BatchDelete
answers 304 with no error; the filer's deletion classifier only forgives
"already deleted" or an exact "not found", so it booked the repeat delete
as a permanent failure. The same mode also compared the fid cookie and
refused chunk manifests with 406, while Go never reads the needle before
those checks when skipping, so the filer's delete of a manifest chunk's
own fid failed permanently too.
The EC branch now reads with read_ec_shard_needle_or_miss and answers as
Go: skipping, a deletion is 304 and any other miss is 500 with Go's text;
checking, every miss is 404 with Go's text ("already deleted",
"locate in local ec volume: FindNeedleFromEcx: needle not found",
"ec shard <vid> not found"). The cookie and manifest checks run only when
the caller asked for the cookie check, which leaves the non-EC path as it
was.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* volume server: match the ReadOnly(VolumeId) variant in write_volume_needles
#11543 matched VolumeError::ReadOnly as a unit variant in Store::write_volume_needles, and #11544 changed it to ReadOnly(VolumeId) in the same merge window. Each passed CI on its own, but master no longer compiles the Rust volume server. Carry the volume id through.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
|
||
|
|
3fd5e4c329 | volume-rust: match the tuple variant VolumeError::ReadOnly in write_volume_needles | ||
|
|
9d3907e36c |
volume: say "volume N is read only" like Go, so filer retries deletes (#11544)
VolumeError::ReadOnly displayed "volume is read-only". Go's store and volume say "volume %d is read only", and the filer's deletion classifier requeues a failed delete only when the error contains "is read only". Against a Rust volume server a BatchDelete on a read-only volume (tier move, maintenance) was booked as a permanent failure and the chunk was never deleted. ReadOnly now carries the volume id and displays Go's text. The text reaches clients through BatchDelete results, the HTTP write and delete error bodies, and gRPC statuses; the gRPC code (FailedPrecondition) and the HTTP/BatchDelete status codes are unchanged. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
bd953b0f84 |
volume server: group-commit fsync writes in the write queue (#11543)
* volume: split the write path into reusable steps do_write_request ran its pre-append checks, the append, the sync rollback, the index publish and the post-write bookkeeping inline, so a batched write could only reuse it one needle at a time. Pull the steps out (check_writable, prepare_write, undo_unsynced_append, publish_write, finish_write) and the store's volume lookup plus disk-space check (writable_volume_mut). do_write_request composes them in the same order with the same early returns; no behaviour change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: group-commit fsync writes in the write queue The write queue holds one store lock for a batch of up to 128 needles but wrote them one at a time, so every fsync needle paid its own .dat sync and its own .idx sync: 2N syncs per batch. Add Volume::write_needles_grouped, after Go's processBatch. A volume's entries are split into runs of distinct needle ids (a repeated id starts a new run, so its dedup and cookie checks see the earlier write). A run with a durable entry appends everything with append_at_ns chained through a local, syncs the .dat once, and only then publishes the entries and syncs the .idx once. A failed .dat sync truncates the .dat back to the run start (marking the volume unavailable if that fails), leaves last_append_at_ns and last_modified untouched, and fails every entry of the run. Runs with no durable entry go through the unchanged per-needle path. Store::write_volume_needles is the queue's entry point; the handlers' non-queue path is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: fail closed when a failed append's rollback fails append_needle discarded the truncate-back result, so a partial write that could not be rolled back left unindexed bytes on the .dat while the volume stayed writable; the next append would bury them mid-file, past the load-time tail check. Route the rollback through undo_unsynced_append, which marks the volume unavailable when the truncate fails, so nothing more is appended over an unverified tail. * volume server: keep a grouped run's I/O error streak from later appends A synced run stages every append before any entry finishes, so the success reset in finish_write ran after the failed appends queued behind the last write to land and erased their media-error streak. Sent one at a time, those errors would have counted and quarantined the volume. Skip the reset when an append after the last landed write added to the streak. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: replay a grouped run's I/O error streak in queue order Skipping the run's success reset whenever an append after the last landed write failed kept the errors from before that write as well, so a run like [EIO, EIO, landed, EIO] reached the quarantine count that the same writes one at a time (one error) do not. Mark the streak where each entry is staged and record the run's success at the last landed write's mark: errors before it are cleared, the ones after it still count. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: replay a run's I/O error streak in one atomic step Reads record their outcomes on the tracker without the volume's write lock, so record_success_at's separate load and store could drop an error a read counted in between, or restore errors a read had just cleared. Keep the count and the clear counter in one atomic word and apply the replay with a single fetch_update. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
7944cb4ba2 |
volume server: keep repeated response headers when proxying a read, like Go (#11542)
In readMode=proxy, proxy_request copied the target's response headers with HeaderMap::insert, so a header the target sent more than once (several Set-Cookie, Vary, Link, ...) reached the client with only its last value. Go's proxyReqToTargetServer adds every value with w.Header().Add. Append instead of insert; the Server header is still dropped and status and body handling are unchanged. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
15d3c65e0c |
volume server: take the heartbeat's remaining store reads off the runtime (#11540)
* volume server: collect EC heartbeats and adjust volume max off the runtime The volume pass moved to the blocking pool, but the heartbeat task still called collect_ec_heartbeat and the following EC shard snapshot, and Store::maybe_adjust_volume_max, directly on its tokio worker. maybe_adjust_volume_max runs statvfs on every auto-sized disk and stats the .dat of every writable volume under the store read lock. All of them block the worker on the node-wide RwLock<Store> whenever a writer holds it or is queued, and every task sharing that worker stalls with it. Run the adjustment, on the pulse and after the master changes volume options, and the EC tick's heartbeat plus shard snapshot through off_runtime, like the volume pass. apply_master_volume_options now only reports whether the options changed; the loop adjusts off the runtime. What is collected and sent, and in what order, is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: keep EC shard deltas out of the volume heartbeat's snapshot EC shard deltas are the diff between the store's shards and last_ec_shards, taken when volume_state_notify fires. But the volume tick and the options-changed heartbeat re-took last_ec_shards from the store too, and a volume heartbeat carries no shard list: a mount or unmount that landed while the notify was pending or the volume pass was collecting was absorbed into the baseline and never sent. The master only learned of it at the next EC tick, 17 pulses later. The EC tick likewise built its heartbeat and its baseline under two separate store guards, so a mount between them was lost the same way. A volume heartbeat now only takes out of the baseline the expired EC shards it reports deleted itself, so the next delta does not repeat them. The EC tick, and the initial EC heartbeat, build the full list and the baseline under one read guard, still on the blocking pool. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: take the heartbeat's remaining store reads off the runtime The heartbeat's volume and EC passes and the volume max adjustment already run on the blocking pool, but several reads of the node-wide RwLock<Store> were still taken directly on the heartbeat's tokio worker: the digest report reset before the first heartbeat, the duplicate-UUID directory lookup and the volume options a master response carries, the EC shard list a state notification is diffed against, and the deregistration heartbeat sent on stop and shutdown. The lock is writer-preferring, so with a writer holding or queued for it each of these parks the worker, and every task sharing that worker stalls with it. Run each through off_runtime, which now takes a closure so a pass can carry what it needs from the master's response. The notify branch's volume snapshot and EC read become one blocking pass, still under two guards in the same order. What is collected and sent, and in what order, is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
fea14c01a7 |
volume server: refuse a tier move while compacting, and a commit once tiered (#11539)
* volume server: refuse a tier move while compacting, and a commit once tiered A tier move to remote and a vacuum compaction of the same volume could interleave and leave the volume unreadable: - A compaction committing while the upload ran swapped .dat/.idx under the transfer, which reopens the .dat by path per part. The move then published an object holding the old (or a mixed) layout against the compacted .idx, and with keep_local_dat_file=false deleted the only compacted .dat. - A tier move finishing while the compaction copy ran (or between the copy and the commit) let the commit swap in the compacted .idx while the reload served the pre-compaction remote object through it. The tier move now refuses to start while the volume is compacting, and re-checks the compaction revision under the store write lock before it records the remote file; on a mismatch it deletes the uploaded object and fails with FailedPrecondition, leaving the volume local. Committing a compaction on a volume that has a remote file is refused and its .cpd/.cpx removed, since the reload would read the remote object through the compacted index. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: abort a tier move whose volume was replaced or removed The tier-up bookkeeping looked the volume up by id only and compared the compaction revision. A delete and re-create of the same id during the upload yields a fresh volume at the same revision, so the move recorded the old volume's object on the new one and, without keep_local_dat_file, removed the new .dat. An unmounted volume was skipped and the move reported success, leaving the uploaded object referenced by nothing. Capture the volume instance (its data-file access control Arc, as the scan and read plans do) with the revision, and require both under the store write lock. A replaced volume fails with FailedPrecondition, a missing one with NotFound; either way nothing is recorded and the object is deleted after the lock is released. Go fails in both cases because deleting or unmounting closes the descriptor its copy reads. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
f2498e122a |
volume: read EC shards fully, like Go's ReadAt (#11537)
EcVolumeShard::read_at and the scrub plan's EcLocalShard::read_at were a single pread/seek_read. That may legally return fewer bytes than asked mid-file (FUSE/NFS/CIFS mounts, a signal, very large requests), and an Interrupted error was not retried. Callers treat a short count as end of file or corruption: verify_ec_shards compared a zero tail and reported a parity mismatch, local scrub reported a broken shard, VolumeEcShardRead ended the stream early, and decode/rebuild/local needle reads failed. Add storage::io::read_full_at, which loops until the buffer is full or a read returns 0 and retries Interrupted, so a short count means EOF. Route both shard read_at methods through it, replace the encoder's private read_at_most with it, and reuse it for the Windows read_exact_at loop. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
62481f1673 |
volume server: ec.decode reads the .ecx from the index dir it was copied to (#11536)
* volume server: ec.decode reads the .ecx from the index dir it was copied to VolumeEcShardsCopy writes the .ecx/.ecj into the receiver's -dir.idx, so with a split data/index dir the decode target has no .ecx beside its shards. VolumeEcShardsToVolume sized the .dat from the right .ecx but built the .idx from the data dir, failing with NotFound after the .dat was already published. It now reads .ecx/.ecj from where the EC volume opened them and writes the .idx beside the .dat, where Go leaves it. The live-entry check and the .dat size also ignored deletions recorded only in the .ecj, which Go folds into the .ecx (RebuildEcxFile) first: a fully deleted volume was decoded instead of reported as having no live entries, and deleted tail needles were copied into the .dat. Both now treat journaled ids as deleted, without rewriting the sealed .ecx. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode keeps the decoded volume writable and reads every .ecj The rebuilt .idx copied a journaled tail needle's .ecx row verbatim after the .dat was cut short before it, so the mount saw a row past EOF and marked the decoded volume read-only. Rows of deleted needles the .dat no longer holds are now dropped, and each journaled needle still in the .dat gets one tombstone instead of one per journal entry. VolumeEcShardsCopy appends journals collected from other holders into the idx dir, but the decode read only the .ecj beside the .ecx, which sits in the data dir when this server generated the shards. It now reads both, once, in bounded chunks via the loader EcVolume uses. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: test ec.decode drops a sealed .ecx tail tombstone Covers the other half of the rule added in the previous commit: a tail needle tombstoned in the .ecx itself (Go's RebuildEcxFile) is cut from the .dat, and its row must not reach the rebuilt .idx either. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
b8f074b7d3 |
volume server: VolumeNeedleStatus reads remote EC shards and reports deleted needles like Go (#11535)
* volume server: VolumeNeedleStatus reads remote EC shards and reports deleted needles like Go
For an EC volume the handler read only locally mounted shards, so a node
that did not hold the shard with the needle's bytes answered Internal
"ec shard N not available locally". Go's ReadEcShardNeedle fetches the
interval from a peer or reconstructs it. It also mapped every regular
volume read error, including a tombstone, to NotFound "needle not found",
which fs.verify treats as a missing needle; Go returns ErrorDeleted as a
plain error ("already deleted"), which fs.verify skips.
The EC branch now drops the store guard and uses the distributed EC read
the HTTP GET path uses. Errors map like Go: needle absent -> NotFound
"needle not found <decimal id>", tombstoned (regular or EC .ecx/.ecj) ->
Unknown "already deleted", anything else -> Unknown with the error text.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* volume server: tell EC deletions and vanished volumes apart in VolumeNeedleStatus
The distributed EC reader returned Ok(None) for an absent needle, a needle
a peer reported deleted, and a volume unmounted after the handler's own
existence check. VolumeNeedleStatus answered all three NotFound "needle not
found", which fs.verify -pruneEntries counts as lost data. A reported
deletion was also lost when an earlier interval failed.
The reader now says why it has no needle (EcMiss: NotFound, Deleted,
VolumeNotFound), classifying the local tombstone itself and letting a
reported deletion outrank other interval errors, as Go's ReadEcShardNeedle
does. VolumeNeedleStatus maps Deleted to Unknown "already deleted" and
VolumeNotFound to "volume not found", and drops its separate EC pre-check.
read_ec_shard_needle_distributed keeps its Ok(None) for every miss, so the
other callers are unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
cc1ec48151 |
volume server: collect EC heartbeats and adjust volume max off the runtime (#11532)
The volume pass moved to the blocking pool, but the heartbeat task still called collect_ec_heartbeat and the following EC shard snapshot, and Store::maybe_adjust_volume_max, directly on its tokio worker. maybe_adjust_volume_max runs statvfs on every auto-sized disk and stats the .dat of every writable volume under the store read lock. All of them block the worker on the node-wide RwLock<Store> whenever a writer holds it or is queued, and every task sharing that worker stalls with it. Run the adjustment, on the pulse and after the master changes volume options, and the EC tick's heartbeat plus shard snapshot through off_runtime, like the volume pass. apply_master_volume_options now only reports whether the options changed; the loop adjusts off the runtime. What is collected and sent, and in what order, is unchanged. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5ece8dd63c |
volume server: drop the unused unmount result in VolumeCopy validation, and test-only EC helpers (#11531)
mount_and_reply ignored the Result of store.unmount_volume when a copied replica failed record count validation, tripping unused_must_use. The Err branch is unreachable there: the volume was mounted under the same store write guard, mount_volume refuses an already loaded vid so it is a fresh Volume with is_compacting false, and a compaction claim needs &mut Volume, i.e. the store lock. Ignore the result explicitly with a one-line reason. Store::delete_expired_ec_volumes and Store::remove_ec_volume are called only from test modules (the heartbeat uses the split find_expired_ec_volumes / remove_expired_ec_volumes halves), so mark them #[cfg(test)]. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
0978e7f833 |
vacuum: keep disk-full read-only volumes reclaimable (#11519)
* storage/topology: keep disk-full read-only volumes vacuumable The vacuum sweep skipped every read-only replica, so a volume that went read-only because its disk filled could never reclaim its garbage — the exact situation compaction exists for. The volume server now reports disk_space_low in VacuumVolumeCheckResponse, and the sweep skips a read-only replica only when the flag is clear. An explicit volumeId vacuum is unaffected: it already bypassed the read-only rule. The field takes number 4: 2 and 3 are downstream-allocated for tombstone retention, keeping the wire merge clean. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * storage: measure vacuum free space against live bytes The pre-compaction space check required the current .dat + .idx size free, which includes the garbage being reclaimed — on a nearly full disk that estimate can never fit, so the volume stayed garbage-bound forever. Measure against the estimated compacted output instead: superblock plus live index entries plus live content bytes, with the existing ten percent buffer unchanged. Mirrors the same check in the Rust volume server. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: count per-needle framing in the compacted-size estimate The live-bytes estimate covered each live needle's content and index entry but not its .dat framing (header, checksum, timestamp, padding — ~32 bytes on version 3). For small-needle volumes that is more than the 10% headroom, so a disk with space between the estimate and the real output still ran out mid-compaction. Rust side mirrors the same formula. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * storage: report disk_space_low only when it is the sole read-only cause Review feedback (ihnokim, greptile, devin): a volume read-only for low disk space AND an operator mark or I/O quarantine was still eligible for the automatic sweep, rewriting a copy meant to stay protected. The flag now reports only the benign sole-cause case in both servers. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * topology: fail closed when the read-only lookup misses in the sweep A heartbeat can drop the volume from the DataNode cache between the location-list copy and VacuumVolumeCheck; a lookup error previously skipped the read-only check entirely. Review feedback (coderabbit). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
43fd5b8d82 |
volume: reclaim staged EC shard generations left by the 2PC switch (#11501)
* volume: remove staged EC generation files on teardown and shard delete The 2PC generation switch stages each run as <base>.ecNN.v<N> plus versioned .ecx/.ecj/.vif files. Nothing on the volume server removes them: isEcDataShardFile only recognises the exact .ecNN name, so the staged files are invisible to every bookkeeping pass, and even full_teardown's wipe-all path left them behind. Each re-encode therefore leaks a full shard set per shard-holding disk. RemoveEcGenerationFiles sweeps <base>.ec*.v<N> and <base>.vif.v<N>, optionally keeping generations at or above a threshold; teardown and the reconcile wipe remove every generation, and a per-shard delete removes that shard's staged generations too. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume: delete staged EC generations older than N via VolumeEcShardsDelete After a 2PC generation switch commits, the superseded generation's <base>.*.v<N> files sit on disk with no cleanup path: teardown removes everything, and a per-shard delete only touches the named shards, so the executor had no RPC that reclaims just the staged leftovers. delete_generations_older_than removes staged generation files strictly below the threshold on every disk. Versioned files are never mounted, so nothing is unloaded first; the committed generation and the canonical files are preserved. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: mirror staged EC generation cleanup Parity with the Go volume server: remove_ec_generation_files sweeps <base>.ec*.v<N> and <base>.vif.v<N> staged by the 2PC switch, called by remove_ec_volume_files (which covers both teardown paths) and the new delete_generations_older_than request field; delete_ec_shards removes a shard's staged generations along with the canonical file. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume: match staged generation filenames literally filepath.Glob interprets metacharacters in the collection part of the base name, so a collection like a[bc] could match another volume's staged files (or miss its own). Scan the directory and compare names literally instead, mirroring the Rust read_dir implementation. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: report generation-sweep errors and drop the store lock first - snapshot the location base names under the read lock and run the filesystem sweep after dropping it, so a slow disk cannot stall the store; - record per-entry read_dir errors in remove_ec_generation_files and propagate them from remove_ec_shard_generations instead of flatten() skipping them; - warn when a staged-shard generation fails to delete rather than reporting success with files left behind. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume: fail shard delete when the staged-generation listing fails A transient ReadDir failure fell back to removing canonical shard names only: staged .v<N> files survived while the RPC still reported success, leaving the leak invisible to retrying callers. ENOENT still means the disk simply has no such directory; other listing errors now propagate. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: propagate staged-generation removal failures delete_ec_shards logged remove_ec_shard_generations errors and the RPC returned success while staged .v<N> files remained, diverging from the Go handler which surfaces the failure. The sweep keeps processing the remaining shards, retains the first error, and volume_ec_shards_delete maps it to Status::internal so callers can retry. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: notify state change even when the shard sweep errors delete_ec_shards already deletes and unmounts the shards before returning a staged-generation failure, so returning early skipped volume_state_notify and the master kept routing to them until the next heartbeat. Notify before propagating the error. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
e57f8c4d87 |
volume server: read GET/HEAD needles off the store lock, and only once (#11487)
The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
67691a1eea |
volume server: split volume_copy into phases and type the delete-after-status gate (#11485)
volume_copy was one ~400-line handler, and the rule that an existing local replica is deleted only after the source's ReadVolumeFileStatus succeeded was held by statement order alone. The keep_remote_data=true that the pre-copy delete and the failed-copy rollback must share was kept in sync by a comment pointing from one to the other. The handler is now a ~60-line orchestrator over connect_to_copy_source, SourceVolumeStatus::fetch, delete_existing_replica, plan_copy_destination and a VolumeCopyJob whose run() drives preallocate_dat, transfer_files, finish_copied_files and mount_and_reply, with cleanup_failed_copy on error. delete_existing_replica takes a &SourceVolumeStatus, which only fetch can construct (private field in a child module), so the delete cannot be called before the status RPC. Both deletes go through delete_replica_keep_remote. Pure refactor: call order, status codes and messages, cancellation checks, throttling, progress reports and cleanup are unchanged. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
02353444ac |
fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts (#11483)
* fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts Port of the Go VolumeCopy hardening in #11238 and #11252. - Pick the destination disk before deleting the existing replica, counting the slot that replica holds as free. If no disk qualifies, the healthy replica is kept instead of being deleted. - Read the source's VolumeStatus before and after the copy. When both succeed and the counts did not change, the mounted replica's file and deleted counts must match; on mismatch it is unmounted and its files removed. A failed "before" read skips the check; a failed "after" read fails the copy. * fix(volume-rust): let a departing caller cancel VolumeCopy's post-copy status read Go reads the source's status after the copy with stream.Context(), so the call ends when the caller leaves. The Rust call had no such link: a source that stalled there held the copied, unmounted files after the caller was gone. Race it against the response channel, like the other blocking steps, so the usual error cleanup removes the partial copy. |
||
|
|
68944e83a3 |
volume: typed tier errors so a missing remote object answers NotFound (#11484)
remote_storage/s3_tier.rs returned Result<_, String> from every
transfer (upload_file, download_file, read_range[_blocking],
delete_file[_blocking]) and from the tier runtime helpers. The tier
move handlers could only wrap that in Status::internal, so a .dat whose
remote object is gone was indistinguishable from an I/O failure to
weed shell.
Add TierError { NotFound, Io, RuntimeUnavailable, Aborted }. Each
variant carries the existing message verbatim. NotFound follows the
rules remote_storage/s3.rs already uses: raw 404 status on HEAD,
NoSuchKey code on GET; a bare 404 on GET stays Io. A progress-callback
Err becomes Aborted. VolumeError gains a transparent Tier variant and
From<VolumeError> for Status maps Tier(NotFound) to NotFound; the tier
move handlers go through status_with_context, so their message text is
unchanged. Every other tier failure is still Internal.
The remote needle read path keeps io::Error::other, so its error kind
and vacuum's handling of it do not change.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
00310f6588 |
volume server: run the vacuum compaction copy without the store lock (#11482)
* volume server: run the vacuum compaction copy without the store lock VacuumVolumeCompact held the store write lock for the whole live-needle copy, including every progress blocking_send on the 16-deep stream. On a large volume that is minutes with every read, write and heartbeat on the node parked behind it, long enough for the master to unregister the node. Split compaction the way Go's CompactByIndex runs it. A short locked step claims the volume's compacting flag, records the makeup_diff watermark (index size and compaction revision) and opens fresh .dat/.idx handles. The copy then replays .idx up to the watermark and copies from those handles with the store lock released; writes that land meanwhile are replayed by makeup_diff at commit, as before. The flag is an Arc<AtomicBool> released when the job is dropped, so every exit path clears it. Because the flag is now visible to other callers, the operations that would pull the files out from under the copy refuse while it is set: unmount (and VolumeConfigure, which unmounts and remounts), delete (checked before the volume is removed from the map, which a refused destroy used to leave unmounted), cleanup, and index relocation. A second compact and a commit stay no-ops, as in Go. The pre-copy fsync is dropped: the copy reads its own handles through the page cache and .cpd/.cpx are fsynced before commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: keep a read-only in-memory index's size for the compaction copy The unlocked copy replays .idx up to index_file_size(). A read-only volume whose .sdx could not be built loads its index into memory without a writer, so that size stayed 0: the copy came out empty and the commit replaced the volume with it. CompactNeedleMap::load_from_idx now records the rows it loaded, which is also what Go's IndexFileSize reports for a read-only index. The copy's index replay now stops reading at the recorded size instead of walking rows appended since, which makeup_diff replays anyway. Adds tests for compacting a read-only volume on both the sorted index and the in-memory fallback, and for VolumeConfigure stopping when the unmount is refused during a copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a vacuum copy as soon as its client is gone The progress callback only noticed a closed response stream when a report was due, every 128 MiB. With the copy now running outside the store lock, a copy nobody waits for keeps the volume marked compacting and so keeps refusing unmount, delete and cleanup until that next report. Check the stream on every callback. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5c9c424a84 |
volume server: stream ReadAllNeedles without holding the store lock (#11481)
* volume server: stream ReadAllNeedles without holding the store lock read_all_needles held store.read() while Volume::read_all_needles read every live needle of the volume into a Vec, and kept holding it through the whole blocking_send loop. Memory grew with the volume, and a slow client parked the scan in a send with the guard held; needle writes and the heartbeat take store.write() on a writer-preferring lock, so the node stopped serving until the client caught up. Take a DatScanPlan (fresh .dat open, end bound) under a short guard and walk it with the guard released, sending one needle at a time. Each record is checked against the live needle map under a brief read guard, as the scan reaches it, and only a live record is parsed, so a damaged stale copy does not fail the stream. Records appended while a pass ran are walked by a follow-up plan, so a needle overwritten during the scan is streamed once, as its new copy. A vacuum commit or re-create of the volume during the scan fails the stream, since the map's offsets no longer describe the pinned file; the plan carries the volume instance and compaction revision for that check. DatScanPlan::scan_records yields records unparsed; scan keeps its behaviour on top of it. Volume::read_all_needles has no caller left and is removed; its tests move to the RPC. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: check ReadAllNeedles liveness only once the send can proceed The per-record liveness check ran before blocking_send, so a scan parked on a full channel held a record it had already judged live. An overwrite landing during that park left the old copy in the stream, and the continuation over appended records then streamed the new copy as well. Reserve channel space first, then take the store read guard, check the record against the needle map and enqueue it through the permit before releasing the guard. The wait for space still happens without the lock; the record is parsed before the guard is taken, and its parse error only counts if the record turns out to be live. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5218e68554 |
volume server: collect heartbeats under the store read lock, off the runtime (#11480)
* volume server: collect heartbeats under the store read lock, off the runtime Every pulse tick, options change and volume-state notification took store.write() for a whole heartbeat pass, directly on the async heartbeat task. The pass fstats every volume's .dat twice and hashes its report, so on a server with many volumes it held the store exclusively for the whole scan: reads and writes stalled, and with the writer-preferring RwLock a pending pass parked every new reader too. The pass only needs to mutate the store for a few rare actions: removing expired EC volumes, deleting expired volumes past their removal delay, and setting no-write on IO-quarantined volumes. It now runs under store.read(), records those as (disk, volume id) actions, and applies them afterwards under a short store.write() that is only taken when there is something to do. Each action re-checks its target under the write lock, so a volume written to, replaced or removed in between is left alone. Expired EC volumes are still removed before the volume pass, as before, because the EC shard count feeds the disk-space-low max volume count. Every pass runs on the blocking pool via spawn_blocking. The heartbeat message is unchanged for the same store state. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: take has_no_ec_shards with the heartbeat's volume list The heartbeat pass took has_no_ec_shards from the EC phase's read lock, then built the volume list under a second one. An EC shard mounted in between went out as "no EC shards" beside a volume list taken after the mount, and the master clears a server's EC registrations on that flag. has_no_ec_shards is now computed under the same read lock as the volume list, with the EC phase's filter: not expired, not quarantined, at least one shard. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
3c1e8ca7a8 |
volume: never finish serving a needle whose data fails its CRC (#11467)
* storage: hold back last chunk until CRC verifies on whole-needle reads Above PagedReadLimit the needle is streamed: headers and body go out before the checksum is computed, so a corrupted needle was served as 200 with bad bytes and readers could not fall back to a replica. The final chunk is now written only after the checksum verifies; on a mismatch the response ends short of Content-Length and the client sees a failed transfer. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: verify needle checksum on streamed reads (parity) Mirror the Go fix: carry the needle checksum in NeedleStreamInfo and have StreamingBody accumulate the CRC and verify it before emitting the last frame; a mismatch ends the body with an error so the client sees the transfer fail rather than receiving corrupt bytes that look complete. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * server: abort the transfer when a committed stream fails A writeFn failure after any byte or WriteHeader call leaves the declared status and Content-Length already sent; http.Error's text then joins the body and can exactly fill the withheld tail of a corrupted needle read — the client sees a complete 200 instead of a failed transfer to retry. Track whether the response is committed (headers sent, or bytes buffered for the deferred flush) and panic with http.ErrAbortHandler instead of appending an error body; pre-commit failures keep the 500 path. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * server: drop the response writer wrapper from the committed-response check Counting buffered writes is enough: with no bytes buffered the status and headers cannot have gone out, and the range branches commit via the explicit WriteHeader call before writeFn runs. The extra ResponseWriter wrapper added a new Write sink site that CodeQL flags. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
0f3ba98e11 |
volume: make volume.scrub report a live needle whose stored id is damaged (#11468)
* storage: scrub live needles' stored id against the index key scrubVolumeData only compared the needle's stored id for tombstones, so header damage on a live needle — where the data CRC cannot see it — passed every scrub mode while reads of that needle kept failing or serving the wrong key's data. Compare the id for every indexed needle. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: scrub live needles' stored id against the index key (parity) Mirror the Go scrub fix: compare the stored needle id with the index key for live needles too, not only for deleted ones. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: cover damaged live needle id in scrub test The tombstone test proved the index-key check fires for deleted entries; add the live-needle mirror of Go's TestScrubVolumeDataChecksLiveNeedleId so a regression in the live path is caught in Rust too. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
8ad2f29e3e |
shell: let volume.deleteEmpty drop volumes with no live needles (#11437)
* shell: let volume.deleteEmpty drop volumes with no live needles The candidate check only accepted a .dat at superblock size, so a volume whose every needle was deleted still had to be vacuumed first — minutes of compaction to rewrite bytes that were all garbage anyway. FileCount counts every indexed entry and DeleteCount every entry made garbage by overwrite or delete, so FileCount <= DeleteCount means nothing live remains and the volume can be unlinked directly. The quietFor guard is unchanged. * volume server: add only_garbage VolumeDelete guard VolumeDelete(only_empty) refuses every volume that ever held data, so a volume whose needles are all deleted could only be removed after a vacuum rewrote it. The new only_garbage flag deletes only when the byte counters show nothing live: DeletedSize covering all of ContentSize, the same all-garbage state vacuum measures. Byte counters are used because the file/delete counts drift on index reload. * rust volume: mirror only_garbage VolumeDelete guard Same check as the Go server: a volume deletes under only_garbage when its deleted bytes cover all content bytes. The grpc handler rejects before the store drops the volume from its map, since destroy errors after removal would still unmount it. * volume delete: let either enabled check pass, keep onlyEmpty on the wire An upgraded shell sending only_garbage to a pre-upgrade server would be read as an unconditional delete (field ignored, only_empty false). The request now keeps only_empty set so old servers check emptiness and refuse, while new servers delete when either check passes. * volume.deleteEmpty: skip remote-backed and protected read-only volumes A remote-tiered replica shares its cloud object with the other replicas, so keepRemoteData=false on one delete removes data they still reference. Protected read-only volumes are quarantined or under maintenance, which is exactly when a replica should not be dropped. * volume delete: validate guarded copies across disks before deleting * volume delete: hold copy locks across guarded validate-and-delete CheckVolumeDeletable released each copy's locks before Destroy ran, so a write landing on a later copy between the two passes refused its destroy after earlier copies were already removed. Pin every copy's dataFileAccessLock (and its location's volumesLock) across validation and removal so a refused delete leaves all copies intact. * volume delete: send deleted-volume notices after releasing locks A blocking send on a full DeletedVolumesChan under volumesLock can stall the heartbeat loop that drains it while it waits on the same locks. Collect the notices under the lock span and send after release. * pb: restore generated-file cosmetics to match the repo's protoc version Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
f7680cf812 |
volume: one long-lived tokio runtime for blocking tiered S3 reads (#11448)
* volume: one long-lived tokio runtime for blocking tiered S3 reads block_on_tier_future, behind read_range_blocking and delete_file_blocking, spawned an OS thread and built a fresh current-thread tokio runtime on every call, then tore the runtime down. On an S3-tiered volume that is once per needle read, per streamed 64 KiB chunk, per DatScanPlan record and per destroy. The SDK client's pooled HTTPS connections are driven by tasks on the runtime a request ran on, so each teardown dropped the pool and every call re-dialed and re-handshook TLS. A panic inside the SDK was also flattened to the fixed string "tier runtime thread panicked". Now one process-wide runtime (OnceLock, multi_thread, 2 workers named tier-io) drives all tier I/O; block_on_tier_future spawns onto it and parks the caller on an mpsc channel for the JoinHandle result. Blocking the caller is unavoidable (the storage layer is synchronous) and is what the old code did through thread::spawn().join(). Handle::block_on is not used because the wrappers are also reached from inside another runtime's worker, where it panics with "Cannot start a runtime from within a runtime". JoinError panics are downcast to &str/String and the payload is kept in the error. Tests cover runtime reuse (Handle::id equal across calls, thread name tier-io), calls from a std thread, from spawn_blocking, and directly from current-thread and multi-thread runtime contexts, and the panic payload. Against the old body 7 of 9 fail. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: return a tier runtime build failure instead of panicking Review follow-up. tier_runtime() expect'ed the runtime build, so an OS refusing threads panicked inside Volume::destroy (after the volume left the in-memory map, before its files were removed) and inside needle reads, bypassing their error paths. Keep the runtime in a Mutex<Option<Runtime>> behind tier_handle() -> Result<Handle, String>: a failed build is returned to the caller through block_on_tier_future's existing Result and is not cached, so a later call retries once the pressure is gone. The lock is held only while building. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: trim comments on the shared tier I/O runtime Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
4299fdf578 |
volume server: VolumeEcShardsDelete full teardown unloads every disk and keeps the shard gauge honest (#11446)
* volume server: VolumeEcShardsDelete full teardown unloads every disk and keeps the shard gauge honest
Go's VolumeEcShardsDelete full teardown calls vs.store.UnloadEcVolume in
the blanket path (weed/server/volume_grpc_erasure_coding.go:488) and
location.UnloadEcVolume in the generation-fenced path (:511): each disk
that had the volume registered drops it, closes its shard descriptors
and gives back its ec_shards gauge before the artifacts are unlinked.
The Rust handler used Store::remove_ec_volume / DiskLocation::remove_ec_volume
instead, which only remove the map entry. Store::remove_ec_volume also
stops at the FIRST disk holding the vid, so on a split-disk volume
(shards on several disks) the blanket teardown left the sibling disks'
EcVolume registered with open fds while the unlink loop deleted their
files underneath it: the heartbeat kept advertising shards whose files
were gone, the inodes stayed pinned by the open descriptors, and the
VOLUME_GAUGE{collection,"ec_shards"} never came back down. The fenced
path leaked the gauge and the descriptors the same way on the one disk
it wiped.
Both paths now use the unload_ec_volume helpers from #11413 (every disk
for the blanket teardown, the strictly-older disk for the fenced one),
and the two Status::internal messages name the disk directory like Go's
"... on %s: %w".
Regression tests build a two-disk store with the same vid mounted on
each disk (the SplitDiskEcFixture, which gains a collection knob so the
gauge read is isolated from parallel tests mounting under "") and assert
that a blanket teardown leaves no EcVolume registered on any disk and
returns the gauge to its pre-mount value, and that a fenced teardown
decrements the gauge for the older disk's shard while preserving the
newer disk. Both fail against the previous handler.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* volume server: trim comments on the EC full-teardown unload path
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
||
|
|
3e9fc9e75b |
volume: reject invalid -minFreeSpace and duration flags at startup like Go (#11447)
* volume: reject invalid -minFreeSpace and duration flags at startup like Go Symptom: the Rust volume server swallowed invalid values for three startup flags and ran with a limit the operator did not ask for. -minFreeSpace=150 became 150 bytes (Go: fatal, percent must be 0..=100), -minFreeSpace=10GiBx and -minFreeSpace=50B fell through to the 1% default (Go: fatal), and -inflightUploadDataTimeout=30sec, =abc or =30 silently became the 60 s default (Go: flag.Duration usage error, exit 2). Go's byte-size units (1,024MB, 42 mib, 100Ki, 1eb) were also not understood and mapped to 1%. Go reference: weed/util/minfreespace.go:73-106 (MustParseMinFreeSpace / ParseMinFreeSpace), weed/util/bytes.go:212-243 (ParseBytes and bytesSizeTable), weed/command/volume.go:226-227 (flag.Duration, i.e. time.ParseDuration grammar). Fix: the two inflight timeout Cli fields are now std::time::Duration with a clap value_parser that ports time.ParseDuration (units ns/us/µs/ μs/ms/s/m/h, fractions, bare 0, int64-nanosecond overflow), so clap prints a usage error naming the value and exits 2 at the same layer as Go's flag package. parse_min_free_spaces returns Result and is built from Go-parity parse_min_free_space / parse_bytes; resolve_config_with_env aborts with Go's "The value specified in -minFreeSpace not a valid value" message, the fail-fast pattern it already uses for --max, --index and --readMode. No new dependency. Negative durations and NaN percent are rejected explicitly (Go quirks with no sane Rust representation). Tests cover the Go fixtures from minfreespace_test.go and bytes_test.go, the duration grammar, and flag-level rejection through Cli::try_parse_from and resolve_config_with_env; the flag-level tests were confirmed to fail against the previous parser. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: say that -minFreeSpace commas separate directories, never thousands Review follow-up. parse_bytes keeps Go's ParseBytes thousands-comma support, but the flag is split on commas into per-directory entries before any entry is parsed (Go's MustParseMinFreeSpace does the same), so -minFreeSpace=1,024MB is the two entries 1 and 024MB. The doc comment and the single-entry tests advertised the comma form as if it were reachable from the flag. Document the split and add a flag-level test: two entries from parse_min_free_spaces, and resolve_config aborting on the directory-count mismatch with one -dir. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: trim comments on the strict flag parsing Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
83d44be0f3 |
volume: detect S3 not-found by typed SDK errors, not the "service error" string (#11444)
* volume: detect S3 not-found by typed SDK errors, not the "service error" string
remote_storage/s3.rs decided ObjectNotFound by matching the Display
output of an aws_sdk_s3 SdkError against "NoSuchKey" / "404" /
"NotFound". In the locked SDK (aws-smithy-runtime-api 1.11.6,
src/client/result.rs:487-497) that Display is a fixed string per
variant, "service error" for every S3 error, so ObjectNotFound was
unreachable: every missing remote object surfaced as
Other("s3 get object: service error") with the real cause discarded.
Go (weed/remote_storage/s3/s3_storage_client.go) uses typed checks:
HEAD (373-374): awserr.RequestFailure with StatusCode() == 404;
GET (436-437): awserr.Error with Code() == s3.ErrCodeNoSuchKey.
read_file now matches SdkError::ServiceError whose GetObjectError
is_no_such_key(); a bare 404 on GET stays a generic error, as in Go.
stat_file matches HeadObjectError::is_not_found() or a raw HTTP 404
status, Go's actual condition. Non-service errors fall through to
Other unchanged. Every SdkError message in s3.rs and s3_tier.rs is
formatted with DisplayErrorContext so the S3 error code and message
survive instead of "service error".
Six network-free unit tests drive the client through a canned
HttpClient (404 NoSuchKey, bare 404 on GET and HEAD, 404 with a
foreign body on HEAD, 403 AccessDenied on GET and HEAD). They need
aws-smithy-runtime-api as a dev-dependency; it is already in the lock
at a single version, so no new crates.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* volume: HEAD not-found is the raw 404 status alone, as in Go
Review follow-up. The HEAD arm also accepted the SDK's NotFound error
code on any status, so a 400 carrying <Code>NotFound</Code> became a
missing object. Go's stat looks only at RequestFailure.StatusCode() ==
404 (weed/remote_storage/s3/s3_storage_client.go:373); do the same. The
raw status still covers the body-less 404 the SDK turns into NotFound
and a 404 whose body names a foreign code. Regression test for the
non-404 NotFound body, which failed against the previous arm.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* volume: trim comments on the typed S3 not-found checks
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: franchb <ru.balloo@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
||
|
|
df4995b894 |
volume server: VolumeMarkReadonly answers NotFound when the volume vanished under the lock (#11443)
* volume server: VolumeMarkReadonly answers NotFound when the volume vanished under the lock
make_volume_readonly looked the volume up, notified the master (step 1),
then took the store write lock (step 2) and marked the volume only `if
let Some(..)`. When the volume left the store during step 1 -- a master
round trip, during which an unmount or a heartbeat expiry can land --
the missing else meant the RPC reported success for a volume the server
no longer has, and step 3 told the master again that it is read-only.
Go's Store.MarkVolumeReadonly (weed/storage/store.go) returns
"volume %d not found" when findVolume comes back nil, and
makeVolumeReadonly (weed/server/volume_grpc_admin.go) returns that error
before the step-3 notification. The Rust step 2 now does the same:
find_volume_mut(vid) -> Status::not_found("volume {vid} not found"), and
the `?` skips step 3, as it already did for a set_read_only_persist
failure. The scrub caller already matches NotFound to skip such a
volume instead of failing the whole report; it now actually gets it.
volume_mark_writable already returns NotFound under its write lock.
The regression test opens the step-1 window deterministically: step 1
awaits the current_master_url read lock, so the test holds its write
guard, lets make_volume_readonly park there after its own lookup
succeeded, unmounts the volume, then releases the guard. With no master
configured the notification is a no-op, so the write lock in step 2 is
the only place left that can notice the volume is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* volume server: trim comments on the vanished-volume mark-readonly path
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
||
|
|
317e756b9a |
volume server: validate ext and collection in gRPC CopyFile/ReceiveFile (Rust) (#11451)
* volume server: validate ext and collection in gRPC CopyFile Port the Go-side checks (checkVolumeFileExtension, checkVolumeCollection) to the Rust volume server so a client-supplied collection or ext carrying a separator or ".." cannot fold a path outside the volume directory. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: validate ext and collection in gRPC ReceiveFile Same port on the write path: the file ReceiveFile creates is built from client-supplied fields, so reject traversal there too. Reported through the response error field, matching Go's SendAndClose. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
bb9942c646 |
volume server: sweep stale EC artifacts before VolumeEcShardsGenerate re-encodes (#11413)
* volume server: sweep stale EC artifacts before VolumeEcShardsGenerate re-encodes The Rust VolumeEcShardsGenerate went straight into write_ec_files: no unload of an already-mounted EC volume and no stale-artifact sweep. Only .ec00..ecNN on the encoding disk were truncated, so a retry could mix two encode runs. A stale N.ec03 left on a sibling disk survived, reconcile later mounted it against the new .ecx, and the new .vif made the encode_ts_ns identity guard pass, so reads served old-run bytes at new-run offsets. Mirror Go's VolumeEcShardsGenerate (#9880 / #9953): UnloadEcVolume on every disk, then removeStaleEcArtifacts on every disk location before encoding. remove_ec_volume_files_full_teardown already has removeStaleEcArtifacts' semantics (.ec00..ec31, .ecx/.ecj/.ecsum[.vN] in both the data and idx dirs, .vif only on a shard-only disk; never the source .dat/.idx), so reuse it. Add Store::unload_ec_volume, which unlike remove_ec_volume does not stop at the first disk and closes the descriptors so the unlink frees the inodes. The store write lock covers only unload + sweep, not the encode. The failure arm now also drops the generation-0 .ecsum, as Go's defer does. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume server: wake the heartbeat after VolumeEcShardsGenerate unloads shards The pre-encode unload drops mounted EC shards from memory, but unlike every other unmount path it did not wake the heartbeat, so the master kept routing reads to shards this server no longer serves until the next pulse. Notify once the store lock is released, and before the sweep error propagates: a failed sweep has unloaded the shards too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume server: clean up encode artifacts when the .vif write fails too Go's shouldCleanup defer covers every error before the .vif commits, not just a failed encode. A serialize or write failure on the .vif left the fresh .ecNN/.ecx/.ecsum behind, which the next generate would have to rely on the new sweep to remove. Extract the cleanup and run it on the .vif error paths as well. * volume server: write the EC .vif atomically Go's SaveVolumeInfo writes a temp file, syncs it, and renames it over the target, so a failed write leaves the previous metadata intact and a read-only .vif fails the save. The direct fs::write truncated the file first, so a write or sync failure could leave an empty .vif even after cleanup_encode removed the generated shards. --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
f0afcf904d |
volume: an EC volume needs an .ecx to mount, and a 0-byte stub never outranks a real index (#11415)
* volume: an EC volume needs a non-empty .ecx to mount Two gaps against Go in how the Rust volume server treats the .ecx. EcVolume::new mounted with no index at all. The per-shard VolumeEcShardsMount path picks the disk by shard file alone, so a shard whose .ecx was on no local directory still registered and was advertised to the master; every VolumeEcShardRead then failed with "ecx file not open", and add_shard's 0-byte guard was neutralised because ecx_file_size stayed 0. Go's NewEcVolume returns an error wrapping os.ErrNotExist. EcVolume::new now fails with NotFound, and Store::mount_ec_shard looks up the .ecx owner across all disks first (findEcxIdxDirForVolume) so a shard on a sibling disk of its index still mounts instead of turning into a hard failure. A 0-byte .ecx stub, as left by a failed EC distribute copy, counted as a valid index. Go requires Size() > 0 wherever the file steers a decision: HasEcxFileOnDisk, findEcxIdxDirForVolume, indexEcxOwners (shared by reconcile and mirror), and VolumeEcShardsCopy removes a copied 0-byte .ecx and fails the copy. Mirror each through one is_usable_ecx_file helper. NewEcVolume itself still accepts a lone 0-byte .ecx as a legitimate empty index, but prefers a non-empty copy, local directory first, over a stub in the other directory; the resolution in EcVolume::new now follows the same order. Tests that mounted EC volumes without any .ecx get a real fixture. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: mount_ec_shard tries every disk; reconcile ignores a 0-byte local .ecx mount_ec_shard returned the first disk's error, so an unusable shard copy (a 0-byte .ecNN left by an interrupted move) hid a good copy on the next disk. Like Go's MountEcShards, keep scanning: NotFound means "not this disk", any other failure is collected, and an all-disks-fail error names every disk tried. "No .ecx on any local disk" is now told apart from "shard not on this server". The orphan-shard reconcile took its locally-mirrored fast path whenever a local .ecx existed at all. A 0-byte stub there registered the shards against an empty index while the owner index skipped that same stub. Go gates the fast path on HasEcxFileOnDisk; do the same. ec_local_ecx_path loses its last production caller and becomes test-only. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: match Go's mount error text and skip the owner stat on the owning disk MountEcShards in Go skips the HasEcxFileOnDisk stat when the disk's own directories already hold the .ecx, dedups a shared -dir.idx across locations in findEcxIdxDirForVolume, and reports "load failures" with the same wording. Also drop two issue-number references from comments. --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
c1ccbcda13 |
volume server: the write queue answers uploads with the needle's real ETag (#11414)
With SEAWEED_WRITE_QUEUE=1 every upload came back with ETag "00000000". The upload handler built the needle with Needle::default(), so its checksum was CRC(0), and handed a clone of it to the queue. The CRC was only computed in the write path, on the worker's clone, and WriteResult carries no checksum back, so n.etag() in the handler formatted the zero checksum. The direct path writes through &mut n and was correct. Compute the checksum in the handler while building the needle, the way Go's CreateNeedleFromRequest does, over the same bytes the write path hashes (the stored data, gzipped or not). The ETag and the has-name flag are read before the write, so the needle is moved into the queue instead of cloned, which also drops a full payload copy per queued upload. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
94a68fa9b9 |
volume: walk_index_file keeps row alignment across short reads (#11412)
* volume: walk_index_file keeps row alignment across short reads walk_index_file issued one Read::read per batch and decoded whatever came back. Read::read may legally return a short count that is not a multiple of the 17-byte entry size (FUSE and network filesystems, a BufReader whose capacity is not a multiple of 17). The split entry at the end of the batch was dropped with no carry and the next read started mid-entry, so every later row was decoded from misaligned bytes and fed to the index as a garbage key/offset/size. This function backs every in-memory index load. Go's WalkIndexFile is immune because it reads through io.ReaderAt, which returns a full buffer or an error. Fill the batch buffer until it is full or the reader reports EOF, retrying ErrorKind::Interrupted, and only then decode whole entries. Reads stay batched at ROWS_TO_READ entries. EOF semantics are unchanged and match Go: on io.EOF Go decodes the whole entries in the final buffer, ignores a trailing partial entry and returns nil. A torn final entry is still skipped without an error here. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: trim walk_index_file comments The batch-fill loop and the ShortReader test helper each carried a paragraph where a sentence suffices. --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
b9ad62fc16 |
[Volume] Keep DAT and index state consistent after async batch Sync failure (#11425)
* fix 11400 * persist failed-recovery quarantine and harden rollback - record the unavailable state in a .unavailable marker, fsync it, and re-arm it on load so a restart cannot serve an unverified pair - quarantine the volume so heartbeats stop advertising it - block MarkVolumeWritable while unavailable, rechecked under noWriteLock - fail every request of a failed batch, not only the succeeded ones - restore the needle map and truncate .dat on inline fsync rollback failure - add truncateIndex for the sorted-file needle map - mirror the fail-closed semantics in the Rust volume server * volume: erase rolled-back mappings instead of leaving tombstones A rolled-back batch or failed inline write used Delete() to undo a needle that did not exist beforehand, leaving a tombstoned map entry whose stale offset makes the next write to that needle fail reading a header that no longer exists. Add removeMapping/restoreMapping to the mappers so recovery erases entries that were absent before the batch and reinstates the exact prior offset/size for ones that were, including tombstones. The index row still goes through Delete so a replay forgets the needle. * volume: gate bulk readers on unavailable and fsync the marker's dir - fsync_dir(&self.dir) synced the volume dir's parent, not the dir holding .unavailable; pass the marker path so the create survives a host crash - export UnavailableError and check it in ReadAllNeedles, VolumeTailSender, VolumeIncrementalCopy, and IncrementalBackup so replica-sync paths cannot stream or append data from an unverified .dat/.idx pair; mirror on the Rust side via read_dat_slice, read_all_needles, dat_scan_plan, and the incremental-copy handler * volume: drop issue references from comments near touched code * volume: stop active scans when the volume becomes unavailable The stream entry-point checks ran once per RPC, so a volume quarantined by a failed recovery mid-scan kept serving data. Recheck availability per needle/chunk on the detached read paths: tail scan and heartbeat, read-all, incremental copy, incremental backup writes, and the Rust StreamingBody chunk reads. Rust incremental copy also rejects a quarantined volume before sync_to_disk touches the backend. --------- Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
d848b8ed00 |
rust volume: regression test for makeup_diff replay across a 32 GiB offset boundary (#11410) (#11427)
* rust volume: test makeup_diff replay across a 32 GiB offset boundary Issue #11410 corrupted a replayed write's index offset in Go's makeupDiff by patching only four of the five offset bytes. The Rust makeup_diff already encodes the whole offset through idx_entry_to_bytes and Offset::from_actual_offset; this adds the mirror of TestConcurrentWriteCrossesOffsetBoundary so a regression would fail here the same way it does under -tags=5BytesOffset on the Go side. Sparse-truncate the .dat to 64 GiB, compact, write, commit: the index offset must equal the .cpd size and the needle must stay readable through a second vacuum. Gated on the 5bytes feature since a 64 GiB .dat exceeds the 32 GiB range of 4-byte offsets. * rust volume: skip the offset-boundary replay test on Windows Windows set_len allocates the full 64 GiB extension instead of a sparse range, so the test fails with StorageFull on CI runners. Gate it to unix, where set_len leaves the extension unallocated. |
||
|
|
0f2ecb766f |
volume server: reject non-ASCII input instead of panicking (#11406)
* volume server: reject non-ASCII input instead of panicking
Three parsers sliced attacker-supplied strings by byte offset, so a
multi-byte character split inside itself and panicked the task:
- parse_needle_id_cookie took the last 8 bytes as the cookie and the
rest as the needle id. Reachable from VolumeServer.BatchDelete,
whose file_ids come straight off the wire as protobuf strings;
that handler already answers 400 per bad fid, so the guard turns a
panicked RPC into the error it was already written to return.
- TTL::read took the unit as the last byte and the count as
everything before it, so "?ttl=5<multi-byte>" split mid-character.
The HTTP upload path does TTL::read(..).ok() and drops an invalid
TTL; AllocateVolume maps the Err to InvalidArgument.
Both now reject non-ASCII up front. Hex and a digits-plus-unit TTL are
ASCII by definition, so no accepted input changes -- covered by tests
alongside the rejection cases.
The six response-* header overrides were inserted with
parse().unwrap(). They come from the query string, so
"?response-cache-control=%0Aevil" decodes to a value HeaderValue
rejects and the unwrap panicked the connection task,
unauthenticated. They now skip the override, matching the if-let the
chunked-response path in the same file already uses.
ReplicaPlacement::from_string was reported as a fourth site but is not
one: reaching chars[2] requires chars[0] and chars[1] to be ASCII
digits, which forces the padded string to be three single-byte
characters, so a multi-byte character always lands on a to_digit()
None first. Kept as a regression test rather than a change.
Each fix was confirmed against the unfixed code first: the parser
tests panic with "byte index N is not a char boundary", and the
integration tests panic at handlers.rs:1413 and ttl.rs:88.
Not a vector, contrary to the report: the HTTP request line. The path
is not percent-decoded before parsing, so "%C3%A9" stays ASCII and
fails the length check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* volume server: fall back to needle MIME when response-content-type is invalid
Skipping an unparseable override left the response without any
Content-Type because the override had already bypassed the normal MIME
selection. Also correct a test comment that described a chars[2] panic
which cannot be reached.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
|
||
|
|
ca62d4297b |
volume: load the .ecj deletion journal in chunks, and repair a torn tail (#11408)
* volume: load the .ecj deletion journal in chunks, and repair a torn tail Two independent defects in the EC deletion journal's load path. 1. The loader issued one NEEDLE_ID_SIZE-byte positional read per entry. That is fine for a healthy journal -- kilobytes -- and pathological for a large one. A `.ecj` is semantically a SET of deleted needle ids but is written as an append-only log that nothing dedupes, and several paths append a peer's ENTIRE journal onto the local one (VolumeEcShardsCopy with copy_ecj_file, EC index recovery, and ec_decode's deliberate cross-holder merge), so a volume whose shards are repeatedly balanced between two servers grows the file without bound. Observed in production: 1.51 TB and 1.30 TB on the two holders of one 10+4 volume containing ~100 distinct ids. At that size the per-entry loop is ~188e9 syscalls, run synchronously while holding the deleted_needles write lock and before the HTTP port opens. The process sits at 100% of one core with a small RSS -- the set stays tiny because the ids repeat -- reading at a few MiB/s because 8-byte reads defeat readahead, logs nothing after "Adding storage location", and ignores SIGTERM. The master then unregisters every volume it holds and reads of them fail. 4.46 and 4.47 are both affected. Read in 1 MiB chunks and build into a local set, merging once at the end so the write lock is not held for the whole scan. Measured on a 256 MiB journal of 100 distinct ids: 33,554,500 syscalls -> 257, identical resulting set. 2. A torn tail silently corrupted later deletes. The journal handle is in append mode, so writes land at the physical end regardless of alignment. A trailing partial record therefore pushed every later append out of alignment: the loader skipped the partial bytes, but the next mount decoded them together with the leading bytes of the following entry, producing one garbage id and dropping the delete that came after the tear -- after acknowledging it. Truncate to a whole number of records at mount, before anything can append. The repair uses its own read+write (non-append) handle: on Windows, append(true) requests FILE_APPEND_DATA without FILE_WRITE_DATA (and .write(true) is subsumed by .append(true)), so SetEndOfFile through the journal handle fails with ERROR_ACCESS_DENIED. The same trap exists in journal_delete's recovery path, which calls set_len on the append handle to roll back a partial write whose sync failed. It is error-handled rather than fatal, so on Windows that rollback silently does not happen. Untouched here; worth a separate fix. Bounding the journal's growth needs compaction, which is deliberately not in this change: replacing the file under a store that can hold several EcVolume instances for one volume id requires coordinating with the other holders, and that belongs at the store layer. Sent separately. Tests: a journal spanning several read chunks loads every entry; a trailing partial record is ignored rather than panicking; a torn tail is truncated at mount and a delete taken afterwards survives a remount. * volume: roll back a failed .ecj append through a dedicated write handle The append handle lacks FILE_WRITE_DATA on Windows, so the set_len rollback after a failed sync silently did nothing and the journal could drift one record past deleted_needles. Same trap as the torn-tail repair in this file; fix it the same way. Also format the new tests. * volume: mirror chunked .ecj load and torn-tail repair in Go --------- Co-authored-by: chrislusf <chrislusf@users.noreply.github.com> Co-authored-by: Devin <devin@cognition.ai> |
||
|
|
4bb40732bb |
volume server: ReceiveFile loses bytes and hides fsync failures (#11407)
* volume server: ReceiveFile loses bytes and hides fsync failures
Three defects in one handler, all on the path that receives a pushed
.dat/.idx/.vif or EC shard:
- `f.write(&content)` never compared the return to content.len().
A short write (ENOSPC, NFS) counted only the bytes that landed,
so every later chunk was written at a shifted offset and the RPC
answered error: "" with a byte count that looked right. Go's
os.File.Write loops. Now write_all.
- `let _ = f.sync_all();` discarded EIO and answered success with
the full byte count. Go omits the check too, but
ReceiveFileResponse carries an `error` field and the caller
renames the staged file into place on success -- so a silent
fsync failure publishes a file whose data never reached the
platter. Flush and fsync failures are now reported.
- Both the per-chunk write and the final fsync were blocking
std::fs calls inside the async fn, on the runtime worker that is
also driving the stream. Switched to tokio::fs + BufWriter, the
shape `drain_copy_stream_to_file` in this same file already uses
and documents. The partial-file cleanup on the error path moves
to tokio::fs::remove_file for the same reason.
The handler had no test at all, which is how the short-write bug
survived. Added a round-trip over a real connection with ragged chunk
boundaries, asserting the bytes on disk and not only the reported
count -- a dropped or reordered chunk changes the file even when
bytes_written still adds up.
That test guards the rewrite; it does not reproduce the original
faults. ENOSPC and EIO need fault injection that this suite has no
harness for, so the short-write and fsync paths are argued from the
code, not demonstrated by a failing test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* volume server: remove the staged file on every ReceiveFile error reply
Flush and fsync failures returned early and left the partial .copying or
shard file behind, as did the pre-existing write-error path. Route all
response-level errors through one cleanup block, matching Go's
close-and-remove on a failed write.
* volume server: tighten ReceiveFile comments
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <devin@cognition.ai>
|
||
|
|
8ff2e0777e |
volume server: HTTP DELETE on a distributed EC volume (#11405)
* volume server: HTTP DELETE on a distributed EC volume
The delete handler validated the cookie with EcVolume::read_ec_shard_needle,
which reads only locally-mounted shards and errors "ec shard N not available
locally" for any interval held by a peer. Every Err was mapped to 500 and no
.ecj tombstone was appended, so on a standard 10+4 spread over 14 servers an
HTTP delete of an EC needle could not succeed. The GET path already goes
through read_ec_shard_needle_distributed.
Route the delete's read through the same distributed reader. It does a
local-first pass in its snapshot phase, so the all-shards-local case costs
what it did before, and no store guard is held across the await (the reader
takes its own; RwLockReadGuard is !Send).
Two smaller corrections fall out of the new return type:
- the reader reports both "needle not in the index" and "volume vanished
between the has_ec check and the snapshot" as Ok(None), which collapses
the old Some(Ok(None)) and None arms into one 404;
- an io::ErrorKind::NotFound now answers 404 rather than 500, matching the
GET path. Telling a caller to retry a delete that can never succeed was
half the bug.
The cookie check and its ordering before the journal append are unchanged.
Not addressed here: Rust journals the tombstone locally while Go routes it to
the primary shard holder. That is a separate behaviour change and belongs in
its own PR against the same issue-10 checkbox.
The regression test mounts 13 of 14 shards, leaving out the one holding the
needle's interval. The distributed reader seeds its Reed-Solomon buffers from
locally mounted siblings, so with >= 10 survivors it reconstructs with no peer
fan-out -- which makes the bug reproducible on a single node. Against the
unfixed handler the test fails with 500 vs 202.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* volume server: fail the delete when the EC volume unmounts mid-request
find_ec_volume_mut returning None used to fall through to a 202 with no
.ecj tombstone written, reporting success for a delete that did not
happen. Answer 404 like the other volume-vanished arms so the caller can
retry after a remount.
* volume server: forward EC needle deletes to a primary-shard holder
Mirror Go's doDeleteNeedleFromAtLeastOneRemoteEcShards: the tombstone is
journaled on one holder of the needle's primary data shard via
VolumeEcBlobDelete (or the local journal when this server holds the
shard), falling back to any other shard holder when the primary has
none. Journaling only on the node that received the DELETE scattered
tombstones across whichever server took the request.
* volume server: route BatchDelete EC deletes through the same forwarding
BatchDelete had the same local-journal divergence as HTTP DELETE, plus a
gap the old code admitted in a comment: the .ecx index cannot supply the
needle's cookie, so EC deletes ran with no cookie check at all. A
distributed read now fills the needle for every EC entry — matching Go's
DeleteEcShardNeedle, which reads and compares the fid cookie even when
skip_cookie_check is set — and the tombstone forwards via
delete_ec_shard_needle_distributed. A needle deleted between read and
journal reports 304 like Go's ErrorDeleted; a vanished volume reports
500 so the filer retries.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <devin@cognition.ai>
|
||
|
|
ac876eef21 |
rust volume: build outgoing gRPC clients in one place and give every dial a connect timeout (#11354)
server/grpc_client.rs stopped at build_grpc_endpoint() -> Endpoint, so all 13 production call sites hand-wrote the same .connect() + X::with_interceptor() + two max_*_message_size() lines. Four of them -- VolumeCopy, VolumeTailReceiver, VolumeEcShardsCopy and the HTTP chunk batch-delete fan-out -- dialed with no timeout at all, so an unreachable peer whose TCP handshake never completes (SYN dropped, blackholed route, host behind a silent firewall) left the operation waiting on the kernel's own retry budget, minutes long. Add GrpcDialOptions (unary / long / stream presets), connect_channel(), and volume_server_client() / master_client() / filer_client() constructors that attach the request-id interceptor and lift both message-size limits, then route all 13 sites through them. build_grpc_endpoint is private again, so connect_channel is the only way out of the module and no call site can dial without picking up a bound. Each site's existing timeouts are preserved exactly; the four bare dials gain a 5 s connect timeout and nothing else. No per-request deadline was added to any streaming call: Endpoint::timeout is a per-request bound on time-to-first-response-headers for every request the channel carries, so a value picked for one short call would also be the header deadline for the whole-volume transfer sharing the dial. The new bound covers the TCP handshake only -- tonic hands connect_timeout to HttpConnector::set_connect_timeout. A peer that completes the handshake and then stalls in the TLS or HTTP/2 exchange is still unbounded at those four sites, as are the RPCs themselves. That is why the three ping_* helpers keep their outer tokio::time::timeout: replacing it with connect_timeout would have narrowed a whole-connect bound they already had. main.rs no longer re-declares GRPC_MAX_MESSAGE_SIZE and the three keepalive/window constants; it imports them from grpc_client.rs so the inbound server and the outgoing clients cannot drift apart. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |