Files
seaweedfs/seaweed-volume/src
d8f926cf46 volume server: stream needle chunks without the store lock (#11488)
* volume server: read GET/HEAD needles off the store lock, and only once

The GET/HEAD handler read the needle synchronously on the tokio worker
while holding store.read(): first a stream-info read that loaded the
whole record just to parse its meta, then, for every needle that was not
streamed (small, compressed, chunk manifest, image ops), a second full
read. For a tiered volume each read is an S3 GET under the store lock,
and a writer queued behind it parks every other store reader.

The regular-volume read now runs in spawn_blocking. Under the store guard
it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat
handle or the remote backend, offset, size); the guard is dropped before
any needle data I/O. No data-file lease is held across the read either,
since a writer waits for one while holding the store write lock. The
index size decides the read, as in Go's readNeedle: a HEAD, a ranged read
or a needle above the stream threshold reads only its header and meta
tail (ReadNeedleMeta) and hands off to StreamingBody or the range path;
everything else is read in full once, with its checksum verified. A
compressed or manifest needle found by the meta read is then read in
full once. The range-from-source read also moves to spawn_blocking.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stream needle chunks without the store lock

StreamingBody::poll_frame took store.read() and find_volume for every
chunk to compare the volume's compaction revision, dup'd the source
handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the
stream also holds a data-file read lease for its whole life, while a
writer waits for that lease under store.write(): the next chunk's
store.read() then waits for the writer and the writer for the stream.

The per-chunk re-lookup was also wrong. The stream reads a handle opened
at plan time, which pins the .dat inode the offset was resolved against;
a vacuum commit renames a new file over .dat and leaves that inode
untouched. The re-looked-up offset belongs to the new file but was read
from the old inode, so a stream whose needle a vacuum moved ended in a
checksum error. The pinned offset stays valid, so the check, and with
it every store access, is dropped, along with the now unused
re_lookup_needle_data_offset and the revision fields of the read plan.

The source is shared as an Arc instead of dup'd per chunk, and the chunk
buffer is a BytesMut that the blocking read hands back with its result,
so its allocation is reclaimed once the previous frame has been written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stop a needle stream once its volume becomes unavailable

Taking the store lock out of StreamingBody also dropped its per-chunk
unavailable_error() check. With -hasSlowRead a writer can take the
data-file lease between chunks, fail its fsync and its truncate, and mark
the volume unavailable; the stream then kept serving the rest of the
needle from its pinned handle.

The volume's io_unavailable reason is now an Arc-shared leaf mutex that
the read plan hands to the stream. Each chunk checks it under its
data-file lease, where the writer marks it, and fails with the same
"volume is unavailable: <reason>" error the old check returned.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-01 23:15:38 +08:00
..