Files
seaweedfs/weed/server/volume_grpc_erasure_coding_recover.go
T
35b090a4df volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go) (#11554)
* volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go)

An EC volume's deletion journal is a set of needle ids, but shard copy
and index recovery appended the peer's whole journal, doubling the file
on every ec_balance round trip. Fold the peer's ids in as a union
instead: only ids the local journal lacks are appended.

- The journal is never replaced. A mounted EcVolume merges a peer's ids
  through its live handle under the lock deletes take (Go
  MergeJournal / Rust merge_journal), wherever its journal lives.
- An unmounted journal gets only the missing ids appended while mounts
  are excluded; the delta is read outside the lock and re-read if the
  journal changed.
- The source .ecj streams into memory as an id set: no staging files,
  chunked reads, memory proportional to distinct ids.
- Go and Rust agree that a source journal exists when it sends a
  modified time or any bytes. A missing source stays a no-op.
- Rust runs every merge in spawn_blocking and shares one receive/merge
  path between shard copy and index recovery.

The decode path and the journal format are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: route .ecj merges to the runtime that holds the journal open

Disks sharing one index directory all resolved as the journal's owner, so
the last one won and a sibling's mounted runtime was skipped: the merge
appended behind its open handle and the sibling kept serving the peer's
deleted needles until remount. Callers now name the receiving disk by its
data directory; the merge goes through that disk's runtime, else a
sibling runtime whose journal is the target file.

In Go the unmounted append now holds every disk's EC lock (in location
order) while it rechecks for a mount, so a sibling mounting from this
disk's index during the unlocked read is merged through instead.

In Rust a mount that lands during the read is merged through directly and
its added count returned, rather than discarded and reported as zero.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: sync merged .ecj records outside the disks' EC locks

The unmounted merge held every disk's EC read lock across its fsync, so a
slow sync on one disk held off mounts on all of them, along with the EC
reads queued behind those mounts. Mounts only need to be excluded while
the records are written: the write now happens under the locks and the
fsync after they are released, since a later mount reads the written
records from the page cache. A failed fsync rolls back only if nothing
has mounted the journal or appended to it since the write.

A merge through a mounted volume now keeps only that volume's disk locked
across its fsync.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: roll back an unsynced .ecj merge through a volume mounted mid-sync

If a volume mounted after the unmounted merge wrote its records but before
the fsync failed, the rollback kept the records because the journal was now
open, leaving ids in the volume's deleted set that may never reach disk; a
retried merge then saw them and synced nothing. The rollback now goes
through that volume the way its own failed journal fsync does: truncate
back and drop the ids from the in-memory set, so a retry appends and syncs
them again. It still keeps the records if the volume journaled since, as
truncating would lose that delete. No fsync runs under the disk locks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: decide .ecj merge rollback from the journal's actual length

Two runtimes can hold one journal (cross-disk mounts). The rollback of an
unsynced merge checked one runtime's cached ecjFileSize, which another
runtime's appends leave stale, so it could truncate a delete that runtime
had already synced. The rollback now holds every holder's journal lock and
truncates only if the file's actual length is still the append's end,
then updates each holder's size and deleted set. Otherwise later records
follow the merged ones, so they stay and are rewritten in place and
synced outside the locks, rather than left possibly not durable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: keep unsynced .ecj merge ids out of mounted deleted sets

When a merge's fsync failed, later records blocked the rollback, and the
rewrite-and-sync failed as well, the merged ids stayed in every mounted
volume's deleted set without being shown durable, so a retried merge saw
them as present and synced nothing. They now leave those sets while the
records stay in the file, matching DeleteNeedleFromEcx, which publishes an
id only after its record syncs. The merge returns the error and a retry
appends and syncs them again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: publish merged .ecj ids to every holder of the journal

Two runtimes can journal into the same file when disks share an index
directory. The merge went through only the first holder, leaving a
sibling's in-memory deleted set without the ids, so it could keep
serving a needle the peer deleted until it remounted. Every holder of
the journal now gets the merged ids, in Go and in the volume server.

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Publish merged .ecj ids to the journal actually written

mountedEcJournal prefers the receiving disk's own runtime for the vid,
whose journal may live in its data directory while the copied records
name a sibling's journal in the index directory. Publishing by the
requested ecjPath then marked a holder of a different file deleted on
records that file never persisted, resurrecting the needles on remount.
Publish by the picked runtime's journal path instead.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:46:36 +08:00

169 lines
6.8 KiB
Go

package weed_server
import (
"context"
"fmt"
"math"
"os"
"time"
"github.com/seaweedfs/seaweedfs/weed/glog"
"github.com/seaweedfs/seaweedfs/weed/operation"
"github.com/seaweedfs/seaweedfs/weed/pb"
"github.com/seaweedfs/seaweedfs/weed/pb/master_pb"
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
"github.com/seaweedfs/seaweedfs/weed/storage"
"github.com/seaweedfs/seaweedfs/weed/util"
)
// ecRecoveryLookupTimeout bounds the master LookupEcVolume call so a slow or
// unresponsive master cannot hang the synchronous recovery RPC.
const ecRecoveryLookupTimeout = time.Minute
// recoverMissingEcIndexes fetches the .ecx / .ecj / .vif index files for EC
// volumes whose shards sit on this server while the index lives only on a peer,
// then mounts the now-recoverable shards. This self-heals the "shards present,
// .ecx missing everywhere local" layout (issue #10104) that the per-disk loader
// and the same-server cross-disk reconcile both leave unmounted, including for
// volumes already broken before the upgrade.
//
// filterVid limits recovery to a single volume id; 0 recovers every orphan on
// this server (used by ec.rebuild to heal volumes the master never learned
// about). It is driven on demand by VolumeEcShardsMount (recover_missing_index),
// so an operator triggers it through ec.rebuild rather than a background loop.
// Returns the number of volumes whose index was recovered.
func (vs *VolumeServer) recoverMissingEcIndexes(filterVid uint32) int {
orphans := vs.store.CollectEcVolumesMissingIndex()
if filterVid != 0 {
filtered := orphans[:0]
for _, m := range orphans {
if uint32(m.VolumeId) == filterVid {
filtered = append(filtered, m)
}
}
orphans = filtered
}
if len(orphans) == 0 {
return 0
}
master := vs.getCurrentMaster()
if master == "" {
glog.Warningf("cannot recover missing EC index without a master connection")
return 0
}
self := pb.NewServerAddress(vs.store.Ip, vs.store.Port, vs.store.GrpcPort)
recovered := 0
for _, m := range orphans {
peers, err := vs.lookupEcVolumePeers(master, uint32(m.VolumeId), self)
if err != nil {
glog.Warningf("ec volume %d: cannot look up peers to recover missing .ecx: %v", m.VolumeId, err)
continue
}
if len(peers) == 0 {
glog.Warningf("ec volume %d: shards present locally but .ecx missing and no peer holds it; leaving shards unloaded", m.VolumeId)
continue
}
if vs.fetchEcIndexFromPeers(peers, m) {
recovered++
}
}
if recovered > 0 {
vs.store.MountRecoveredEcShards()
glog.V(0).Infof("recovered missing EC index for %d volume(s) from peers and mounted their shards", recovered)
}
return recovered
}
// lookupEcVolumePeers asks the master which servers hold shards for vid and
// returns the unique peer addresses, excluding this server itself.
func (vs *VolumeServer) lookupEcVolumePeers(master pb.ServerAddress, vid uint32, self pb.ServerAddress) ([]pb.ServerAddress, error) {
ctx, cancel := context.WithTimeout(context.Background(), ecRecoveryLookupTimeout)
defer cancel()
var peers []pb.ServerAddress
seen := make(map[pb.ServerAddress]bool)
err := operation.WithMasterServerClient(ctx, false, master, vs.grpcDialOption, func(client master_pb.SeaweedClient) error {
resp, err := client.LookupEcVolume(ctx, &master_pb.LookupEcVolumeRequest{VolumeId: vid})
if err != nil {
return err
}
for _, shardIdLocations := range resp.ShardIdLocations {
for _, loc := range shardIdLocations.Locations {
addr := pb.NewServerAddressFromLocation(loc)
if addr.Equals(self) || seen[addr] {
continue
}
seen[addr] = true
peers = append(peers, addr)
}
}
return nil
})
return peers, err
}
// fetchEcIndexFromPeers tries each peer in turn, copying the .ecx (required) and
// the .ecj / .vif (best-effort) for the volume onto the local disk recorded in
// m. The .ecx is an immutable encode-time index, identical on every holder, so
// any peer's copy serves. The .ecj is a per-holder deletion journal that differs
// across holders (a delete is journaled on only one node); the recovered node
// adopts the source peer's deletion view, exactly as a balanced or rebuilt shard
// does — the EC delete model already tolerates that divergence. The first peer
// that yields a non-empty .ecx wins; a peer with no index or a 0-byte stub is
// skipped (the orphan shards are non-empty, so a 0-byte index cannot be theirs).
// A failed copy removes its partial file so a later attempt is not blocked by a
// stub.
func (vs *VolumeServer) fetchEcIndexFromPeers(peers []pb.ServerAddress, m storage.EcVolumeMissingIndex) bool {
idxBaseFileName := storage.VolumeFileName(m.IdxDir, m.Collection, int(m.VolumeId))
dataBaseFileName := storage.VolumeFileName(m.DataDir, m.Collection, int(m.VolumeId))
ecxPath := idxBaseFileName + ".ecx"
removePartial := func(path string) {
if err := os.Remove(path); err != nil && !os.IsNotExist(err) {
glog.Warningf("ec volume %d: remove partial %s: %v", m.VolumeId, path, err)
}
}
for _, peer := range peers {
err := operation.WithVolumeServerClient(true, peer, vs.grpcDialOption, func(client volume_server_pb.VolumeServerClient) error {
// .ecx is mandatory; a peer without it errors and we move on.
if _, err := vs.doCopyFile(client, true, m.Collection, uint32(m.VolumeId), math.MaxUint32, math.MaxInt64, idxBaseFileName, ".ecx", false, false, nil); err != nil {
removePartial(ecxPath)
return err
}
info, statErr := os.Stat(ecxPath)
if statErr != nil {
removePartial(ecxPath)
return fmt.Errorf("stat copied .ecx %s: %w", ecxPath, statErr)
}
if info.IsDir() || info.Size() == 0 {
removePartial(ecxPath)
return fmt.Errorf("peer %s served an unusable .ecx (size %d)", peer, info.Size())
}
// .ecj is the source peer's deletion journal; .vif carries EC params
// and EncodeTsNs. Both are best-effort: a missing .ecj is recreated at
// mount and a missing .vif falls back to default EC parameters. The
// journal is a *set*: merge it as a union so a bounced volume
// cannot double it. The merge only appends whole records, and the
// .vif copy stages and renames, so a failure leaves nothing to clean.
if err := vs.copyEcjAndMerge(client, m.Collection, uint32(m.VolumeId), m.DataDir, idxBaseFileName, util.NewWriteThrottler(vs.maintenanceBytePerSecond)); err != nil {
glog.Warningf("ec volume %d: copy .ecj from %s: %v", m.VolumeId, peer, err)
}
if _, err := vs.doCopyFile(client, true, m.Collection, uint32(m.VolumeId), math.MaxUint32, math.MaxInt64, dataBaseFileName, ".vif", false, true, nil); err != nil {
glog.Warningf("ec volume %d: copy .vif from %s: %v", m.VolumeId, peer, err)
}
return nil
})
if err != nil {
glog.V(1).Infof("ec volume %d: fetch missing .ecx from %s failed: %v", m.VolumeId, peer, err)
continue
}
glog.V(0).Infof("ec volume %d: fetched missing .ecx from %s into %s", m.VolumeId, peer, m.IdxDir)
return true
}
return false
}