mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-06 06:22:05 +02:00
d7a02567e38dba8eeeaa0f309db48afc0bfa17be
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6c07a5fdd0 |
s3: keep small ranged GETs on range reads, no whole-chunk downloads (#11577)
* filer: keep a ranged read in random mode through its contiguous tail A far ReadAt on a fresh ReaderPattern left the sequential counter at -1, so the next buffer of the same ranged request landed on the frontier and flipped the verdict straight back to sequential — readChunkSliceAt then paid a whole-chunk fetch for the remainder of the range. Drop the counter to -ModeChangeLimit when random mode is entered so the verdict needs sustained sequential evidence to undo, matching the hysteresis an established sequential stream already gets. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: pin small ranged GETs to range reads A ranged GET whose first read lands within SeqTolerance of offset 0 is judged sequential immediately, and even a far-starting range could flip back mid-request; either way readChunkSliceAt downloads each covered chunk in full, multiplying disk reads for small ranged reads (measured ~7x). Pin random mode for ranged requests no larger than SeqTolerance so all of the request's buffer reads stay range fetches. Larger ranges keep the dynamic pattern, where whole-chunk fetches amortize. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: fetch only the part of a chunk the view covers Replaces the PinRandomMode size heuristic with a per-chunk coverage rule. ViewFromVisibleIntervals already clips chunk views to the request window, so a view that is not IsFullChunk() is one the request only partially needs; fetch it as a range regardless of the detected read pattern. This closes the holes a request-size pin left open: ranges larger than SeqTolerance no longer revert to whole-chunk downloads once their buffers look sequential, and ranges that fully cover a chunk keep the shared whole-chunk path instead of fetching 256KiB slices piecemeal. Prefetch (MaybeCache) skips clipped views so it cannot amplify a range read either. PinRandomMode is dropped: no caller needs it once coverage drives the fetch choice. Range fetches route through fetchChunkDataFn so tests observe them the same way as whole-chunk downloads. * filer: keep ciphered chunks on the whole-chunk path A range fetch cannot save bytes for a ciphered chunk: readEncryptedUrl always downloads and decrypts the whole blob before slicing. Sending partial views of ciphered chunks through fetchChunkRange would repeat the full download per buffer, so they keep the shared whole-chunk path where one download serves every buffer. Prefetch stays enabled for them for the same reason. * filer: keep compressed chunks on the whole-chunk path Like ciphered chunks, a range request on a compressed chunk makes the volume server read and decompress the whole needle, so range-per-buffer would repeat the full backend read for each 256KiB window. Route them through the shared whole-chunk path via ChunkView.CanRangeFetch. * filer: fall back to range fetch when a chunk exceeds the reader budget A ciphered or compressed chunk larger than readerCacheSizeMB can never be read through the whole-chunk path — the budget rejects the buffer — so its partial views must still range-fetch or the GET fails outright. --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
9b3b12c607 |
filer: pin-aware reader cache eviction and stream release (#11503)
* filer: synchronize stream pins and release them on transitions Guard chunkStream.cacher with the ReaderCache lock everywhere: mount sections share one ChunkReadAt across concurrent reads, and unsynchronized release could double-unpin. Reads served from the chunk cache now detach the stream's pin instead of retaining the previous chunk. Eviction prefers unpinned downloaders so a pinned buffer is not dropped mid-stream. A new ReleaseStream lets callers drop their pin without destroying the shared cache; S3 and WebDAV readers use it. lastChunkFid becomes atomic since concurrent mount reads can update it. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: keep eviction bounded when every downloader is pinned Both eviction paths still fall back to a pinned victim when no unpinned one exists, so abandoned stream pins cannot bypass the downloader limit or stall the memory budget. Budget eviction also rechecks the pin under the ReaderCache lock at removal time: a stream that pinned the selected victim in between keeps it mapped and the selection retries. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: restore budget bookkeeping when a victim gets pinned mid-eviction removeUnpinned losing the pin race left the victim out of the idle list while still holding its reservation, making it unevictable even as the pinned fallback. Push it back when the reservation is still live. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
c72eda50a8 |
s3: drop implicit reader cache budget that throttled S3 GETs (#11384)
* fix(filer): leave reader cache unbounded without an explicit budget NewReaderCache silently installed a 256MiB ReaderCacheBudget when the caller passed none. Only weed mount opts into a budget; every other caller (S3 gateway, WebDAV, query engine, mq logstore) inherited the cap. Under ~90 concurrent S3 GETs of medium objects, prefetch wants far more than 64 chunk buffers, so reserve() serialized chunk fetches, clients timed out and retried, and the retry re-downloaded chunks the cancelled request had already fetched. A nil budget now means unbounded, restoring the pre-4.47 behavior for callers that never asked for a memory cap; reserve/complete/release are nil-safe. The mount path is unchanged and still enforces -readerCacheSizeMB. Fixes #11380 * feat(s3): expose -s3.readerCacheSizeMB reader buffer budget Operators who want the S3 gateway read path memory-bounded can now opt in: -s3.readerCacheSizeMB on weed filer/server/mini and -readerCacheSizeMB on standalone weed s3, matching the mount flag. The default 0 keeps the unbounded pre-4.47 behavior; a positive value installs a shared ReaderCacheBudget across in-flight and retained chunk buffers for all S3 GETs. * fix(filer): validate chunk size before consulting the reader budget A nil budget returned early and skipped the negative chunkSize check, letting a corrupted size reach mem.Allocate and panic. Also drop the command-specific flag prefix from the S3 validation error since standalone weed s3 exposes the option as -readerCacheSizeMB. * filer: drop chunk buffers once fully consumed ReaderCache retained every completed chunk buffer in the downloaders map until the slot limit evicted it, so buffers lingered after all readers finished with them. Track attached readers on each SingleChunkCacher and remove the cacher when the last reader consumes the buffer to its end. In-flight download deduplication and the prefetch handoff are unchanged: a buffer always survives until fully read, partial reads keep it available, and an attached reader pins a consumed buffer until it detaches. Repeat reads now go through the chunk cache where enabled, or refetch. * filer: drop consumed buffers on last detach, rechecked under cache lock Two review findings on the drop-on-consume change: - Removal only fired when the detaching reader itself reached the chunk end. If the end-reaching reader finished first and the last remaining reader did a partial read or cancelled, the consumed buffer and its budget reservation lingered until eviction. Track a persistent consumed flag instead, so any end-reaching read marks the buffer and the last detach drops it. - remove() checked only map identity, so a reader attaching between the reader count hitting zero and removal could attach to a cacher that was then deleted underneath it. removeConsumed() re-checks identity, readers == 0, and consumed under the ReaderCache lock; a raced attach keeps the cacher and its own detach retries the removal. |
||
|
|
4a1d65939f |
fix(mount): bound reader cache memory across open files (#11220)
* fix(filer): bound retained reader cache buffers by bytes * test(filer): keep in-flight downloads during cache trimming * feat(mem): expose pooled allocation capacity for byte reservations * fix(mount): share a configurable reader buffer budget across files * fix(filer): release failed prefetch slots and memory reservations * feat(mount): expose a soft Go runtime memory limit * docs(filer): restore shared-download rationale in startCaching The one-line comment replacing the original context.Background() explanation was too thin for readChunkAt to cross-reference shared resource semantics. Restore a concise note on why request cancellation must not abort a download shared by concurrent readers. * test(filer): loosen reader cache test deadlines to 5s Three tests used 1-second deadlines that can flake on CI under load: TestReaderCacheBudgetInFlight, TestReaderCacheEvictionDoesNotHoldCacheLock, and TestReaderCacheFailedPrefetchReleasesBudget. Increase to 5 seconds. * test(filer): cover re-read after reader cache eviction Add TestReaderCacheReReadAfterEviction: reads chunk 'a', reads chunk 'b' (evicting 'a' via budget pressure), then re-reads 'a' and asserts a fresh download returns correct data. Verifies the core correctness property that eviction never exposes missing or stale data to readers. |