mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-08 15:27:43 +02:00
* rust volume: add a .dat scan plan that runs without the store lock DatScanPlan captures a fresh .dat handle, the version, the start offset and an end bound while the caller holds a store guard, then visits one record at a time with positional reads that never touch the Volume, the way Go's ScanVolumeFileFrom feeds a scanner. The handle pins the inode the offset was resolved against: a vacuum commit renames .cpd over .dat and destroy unlinks it, and neither rewrites the pinned bytes. The end bound is read while no writer can hold store.write(), so the scan never meets a partial append. It is a fresh open, not try_clone, because on Windows read_exact_at uses seek_read, which moves a cursor a clone shares with the writer. A header whose size is negative, or does not fit before the end bound, ends the pass before the body length is computed or anything is allocated. In today's scan a negative size reaches needle_body_length and either overflows the buffer size or walks the scan from a wrong offset. A size near i32::MAX overflows padding_length's i32 arithmetic, which panics in debug builds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3 * rust volume: stream the tail scan with the store lock released volume_tail_sender read every needle from the start offset to EOF into a Vec while holding store.read(). volume.merge tails from zero, so that was the whole volume in memory. And because needle writes and the heartbeat take store.write() on a lock that prefers writers, the whole node stopped serving until the scan finished: the failure #11235 fixed for EC scrub. Each pass now runs on a blocking thread. Under one store guard it resolves the start offset and captures a DatScanPlan, then drops the guard and sends each needle as it is read, as Go's VolumeFileScanner4Tailing does. This replaces the one-guard-across- search-and-scan rule from the previous commit with a stronger invariant: the offset, the handle and the end bound come from the same guard, and the handle pins the inode, so a vacuum commit mid-scan cannot point the offset into the compacted file. A scan error now ends the stream with Status::internal instead of a clean EOF, as Go's `streamFollow: %w` does. Once needles stream, a clean EOF after a partial pass would let volume.move treat a truncated tail as complete. A panic in the pass is reported the same way. A receiver that hangs up is also noticed between skipped needles, not only on a send. Unchanged: the append_at_ns filter, the header on every 2MB chunk, the caught-up heartbeat without a scan, and the draining countdown. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3 * rust volume: fail the tail pass on a short read below the snapshot end DatScanPlan::scan treated an UnexpectedEof on the header or body read as the end of the data and returned Ok. Every byte below the captured end existed when the plan was taken, so a short read there can only mean the inode was truncated under the plan: an unmount followed by a VolumeCopy of the same volume id reopens .dat with truncate(true). The pass then reported Scanned, the next pass found the volume gone, and the stream ended cleanly after a prefix of the planned records, which volume.move would take as a complete tail. Both short-read arms now fail the scan with an I/O error that names the offset and the snapshot end, so tail_pass reports Status::internal as it does for every other read failure. The break arms were carried over from scan_raw_needles_from, where the whole scan ran under the store guard and nothing could truncate the file. Found by the Devin and Greptile reviews on #11275. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * rust volume: sum the needle padding in i64 so a corrupt size cannot overflow padding_length added the header, checksum and timestamp widths to the needle size in i32. A size read from a corrupt header can sit near i32::MAX, and that sum then overflows: a panic with overflow checks, a wrapped padding without. DatScanPlan::scan bounds the size against the bytes left before computing the body length, but that only keeps such a size out of the arithmetic while under 2 GiB of the file remains, so on a large volume the scan could still reach the overflow and, in release, size a buffer from garbage. Sum in i64 in both version branches. The result is at most NEEDLE_PADDING_SIZE, so it still fits Size. The scan comment no longer claims the bound check prevents the overflow. Found by the CodeRabbit review on #11275. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * rust volume: propagate dat scan parse failures --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com>