Files
seaweedfs/weed/server
Chris LuandDevin 14fdd61aea filer: stop the aggregated metadata subscribe loop rescanning an exhausted persisted log (#11644)
* fix(filer): gate the aggregated metadata disk pass on real change

A subscriber whose start position is past the end of the local persisted
log re-ran the whole persisted-log pass - store listings, file opens,
readahead - on every loop iteration. Each iteration is paced only by the
shortest wake (the 20ms hold floor on a busy watermark), so one parked
subscriber kept a full CPU core busy for the life of the stream.

The aggregated loop now mirrors the local loop's gate: the disk pass
runs on the first pass and afterwards only when something it cannot
miss changed - a local flush landed, the peers' flush low-watermark
advanced (more content admitted, or new files in a shared store), the
cursor moved, or a disk hold is pending (the ring read that follows an
empty pass parks internally, so skipping there would strand a held
entry).

Regression test: a subscriber parked past the persisted-log tail holds
the listing rate near zero and still delivers once peers report
progress.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: re-arm the aggregated disk pass on unobserved change

Review found three staleness classes the gate could not see: the flush
low-watermark only catching rises (a joining peer lowers the minimum and
invalidates an earlier pass's proof), a peer past the minimum landing a
file without moving it, and a chunk subscriber's refs-stop bound
advancing with wall time. Re-read when the low-watermark moves in either
direction, when the chunk listing bound admits more files, and on a slow
re-probe cadence for files no watermark can signal.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: unwind the parked ring read so the disk re-probe runs, and re-read on cursor rewinds

A caught-up subscriber parks inside LoopProcessLogData's wait loop, so
the re-probe interval in the outer disk gate could never elapse there;
the callback now unwinds the read once the cadence is due so the gate
re-evaluates. The cursor trigger also needs to notice rewinds, not just
advances, since ResumeFromDiskError moves the cursor backward.

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 22:03:31 +08:00
..
2026-02-20 18:42:00 -08:00