Files
seaweedfs/weed/command
562afa8ec9 filer: resume metadata subscriber from processed watermark on reconnect (#11574)
* filer: resume metadata subscriber from processed watermark on reconnect

* filer: take the reconnect position from GetResumeTsNs verbatim

The callback is the subscriber's durable resume point; falling back to
StartTsNs when it returns zero can resume from a cursor the log-chunk
reader advanced past still-pending work.

* filer: advance the stream cursor once a retried event recovers

RetryForeverOnError resolves the failure inside handleErr, so returning
without moving StartTsNs replays work the event already did when the
stream reconnects before the next one arrives.

* filer: let filtered-progress markers move the processed watermark

A marker means the source examined everything up to its timestamp and
skipped what did not match the subscription. With a resume callback the
marker now reaches the consumer, and AddSyncJob advances the watermark
to it once every earlier job finished and no failure pins the offset.
Idle filtered stretches no longer rescan on every reconnect, while the
guards keep the watermark behind pending or failed work.

* filer: unpin the watermark once a failed event completes

oldestFailedTsNs was only ever set, so a failure that a replay later
fixed still held the resume offset, and every reconnect re-read the
same backlog. Track outstanding failures in a set and recompute the
pin when the failed event's job finally succeeds.

* filer.remote.gateway: resume bucket sync from the processed watermark

The bucket-sync subscriber runs the same MetadataProcessor queue as
filer.remote.sync; give it the same GetResumeTsNs callback so a
reconnect resumes from durably processed work, not the last seen event.

* util: treat a peer-sent gRPC Canceled as transient

A peer tearing down its end of the transport reports codes.Canceled
("the client connection is closing"), which IsTransientError used to
reject: the sync job then failed on the first try and held the offset
until a restart. Caller's own cancels are still excluded up front by
errors.Is(err, context.Canceled), so only teardown-style statuses take
the new branch.

* fix: preserve filtered progress and distinguish caller cancellation

* filer: bound the failed-event ledger past a persistent outage

A destination rejecting every event grew failedTs by one entry per source
event for the life of the processor. Past maxFailedSyncEvents the set now
collapses to a sticky pin at the smallest failure seen, so the watermark
still replays from the oldest failure while memory stays bounded; a
restart re-derives the exact set.

Also keep a resume-callback consumer's chunk-ref replay filter at the
subscribe-time position instead of option.StartTsNs, so a resubscribe does
not filter out events whose async processing is still pending.

* filer: key the failed-event ledger by event, not just timestamp

A success for one event cleared the pin recorded for a different event
that shared its TsNs, letting the watermark pass an unresolved failure.
The ledger now keys on the event's path identity, so recovery unblocks
only the event that actually failed.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 20:57:50 +08:00
..
2026-08-28 20:46:59 -07:00
2026-02-20 18:42:00 -08:00
2025-12-14 16:02:06 -08:00
2026-02-20 18:42:00 -08:00
2025-10-13 18:05:17 -07:00
2019-11-28 18:44:27 -08:00
2026-02-20 18:42:00 -08:00