mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-06 06:22:05 +02:00
* vacuum: bound the commit RPC with a phase deadline VacuumVolumeCommit ran on context.Background(), so a volume server that keeps the call pending would hold the topology-wide vacuum guard forever and every later sweep would be skipped. Give the call a deadline scaled like the existing phase waits (one minute per GB of the volume size limit) so a stalled commit ends as an error instead of blocking the sweep; the timeout is a var so tests can shrink it. * vacuum: bound the replica status probe with a phase deadline The VolumeStatus call on replicas that were not compacted also ran on context.Background(), so a stalled replica could pin the sweep the same way a stalled commit can. Give it the same per-phase deadline. * vacuum: bound the cleanup RPC with a phase deadline VacuumVolumeCleanup also ran on context.Background(); a stalled server would keep the sweep worker and the shared vacuum guard pending forever. Give it the same per-phase deadline. * vacuum: let the check and compact phase waits cancel their RPCs The coordinator wait timers fired while the check and compact calls still ran on context.Background(), so the sweep gave up but the RPC goroutine stayed until the server answered, and a compact stream kept writing on the server. Share one deadline context between the wait and the calls so an expired wait actually cancels them. * vacuum: test that a stalled volume server releases the vacuum guard A fake volume server keeps one vacuum-phase RPC pending until the client context is cancelled. Before the phase deadlines, Vacuum never returned and vacuumLockCounter stayed held; now each phase cancels on its deadline and the guard is free for the next request. * volume: stop compaction at the next needle when the client cancels The progress callback only noticed a gone client when a 128 MiB report failed to send, so an aborted VacuumVolumeCompact kept copying for up to a whole interval while the master had already moved on to cleanup. Check the stream context on every needle, the same early return the Rust volume server does with tx.is_closed(). * vacuum: assert the stalled phase RPC is cancelled, not just bypassed The check and compact coordinator waits already returned on timeout before the deadlines existed, so a regression that put the calls back on context.Background() would pass unnoticed. Wait for the fake server to report that the phase RPC context ended. * vacuum: give the stalled-RPC test room to reach the handler The 50ms phase budget starts before goroutine scheduling and the gRPC dial, so a busy test host could expire it before the fake server saw the call. Raise the override to 250ms; the test still finishes in about a second. * vacuum: describe the phase deadline as scaled, not per-GB The formula keeps the exact expression the check and compact waits already used (floor plus one at 1 GiB granularity); it is a backstop, not a per-GB SLO.