Files
seaweedfs/weed/storage/erasure_coding/ec_teardown.go
T
Chris LuandDevin 43fd5b8d82 volume: reclaim staged EC shard generations left by the 2PC switch (#11501)
* volume: remove staged EC generation files on teardown and shard delete

The 2PC generation switch stages each run as <base>.ecNN.v<N> plus
versioned .ecx/.ecj/.vif files. Nothing on the volume server removes
them: isEcDataShardFile only recognises the exact .ecNN name, so the
staged files are invisible to every bookkeeping pass, and even
full_teardown's wipe-all path left them behind. Each re-encode therefore
leaks a full shard set per shard-holding disk.

RemoveEcGenerationFiles sweeps <base>.ec*.v<N> and <base>.vif.v<N>,
optionally keeping generations at or above a threshold; teardown and the
reconcile wipe remove every generation, and a per-shard delete removes
that shard's staged generations too.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: delete staged EC generations older than N via VolumeEcShardsDelete

After a 2PC generation switch commits, the superseded generation's
<base>.*.v<N> files sit on disk with no cleanup path: teardown removes
everything, and a per-shard delete only touches the named shards, so the
executor had no RPC that reclaims just the staged leftovers.

delete_generations_older_than removes staged generation files strictly
below the threshold on every disk. Versioned files are never mounted, so
nothing is unloaded first; the committed generation and the canonical
files are preserved.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: mirror staged EC generation cleanup

Parity with the Go volume server: remove_ec_generation_files sweeps
<base>.ec*.v<N> and <base>.vif.v<N> staged by the 2PC switch, called by
remove_ec_volume_files (which covers both teardown paths) and the new
delete_generations_older_than request field; delete_ec_shards removes a
shard's staged generations along with the canonical file.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: match staged generation filenames literally

filepath.Glob interprets metacharacters in the collection part of the
base name, so a collection like a[bc] could match another volume's
staged files (or miss its own). Scan the directory and compare names
literally instead, mirroring the Rust read_dir implementation.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: report generation-sweep errors and drop the store lock first

- snapshot the location base names under the read lock and run the
  filesystem sweep after dropping it, so a slow disk cannot stall the
  store;
- record per-entry read_dir errors in remove_ec_generation_files and
  propagate them from remove_ec_shard_generations instead of flatten()
  skipping them;
- warn when a staged-shard generation fails to delete rather than
  reporting success with files left behind.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: fail shard delete when the staged-generation listing fails

A transient ReadDir failure fell back to removing canonical shard names
only: staged .v<N> files survived while the RPC still reported success,
leaving the leak invisible to retrying callers. ENOENT still means the
disk simply has no such directory; other listing errors now propagate.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: propagate staged-generation removal failures

delete_ec_shards logged remove_ec_shard_generations errors and the RPC
returned success while staged .v<N> files remained, diverging from the
Go handler which surfaces the failure. The sweep keeps processing the
remaining shards, retains the first error, and volume_ec_shards_delete
maps it to Status::internal so callers can retry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: notify state change even when the shard sweep errors

delete_ec_shards already deletes and unmounts the shards before
returning a staged-generation failure, so returning early skipped
volume_state_notify and the master kept routing to them until the next
heartbeat. Notify before propagating the error.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:55:25 +08:00

135 lines
4.8 KiB
Go

package erasure_coding
import (
"context"
"errors"
"fmt"
"os"
"path/filepath"
"strconv"
"strings"
"github.com/seaweedfs/seaweedfs/weed/operation"
"github.com/seaweedfs/seaweedfs/weed/pb"
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
"google.golang.org/grpc"
)
// ErrFullTeardownNotAcked marks a reachable server that completed the delete
// RPC but did not report a full teardown (e.g. a pre-upgrade volume server), so
// a stale EC generation may remain on it. Callers distinguish this from an
// unreachable node (which may recover and be re-swept) with errors.Is.
var ErrFullTeardownNotAcked = errors.New("delete did not perform full teardown (pre-upgrade volume server?); a stale EC generation may remain")
// UnmountAndDeleteEcShards unmounts then tears down the named EC shards for a
// volume on one server. Unmount must precede delete (delete requires the shard
// be unmounted); both RPCs are idempotent against missing shards.
//
// encodeTsNs fences both RPCs:
// - 0 selects the server's blanket, generation-independent teardown. This is
// the correct choice for a pre-encode or rollback wipe: it clears same-
// generation shards (a retried encode's prior attempt shares the job's
// generation) and shards whose .vif generation is unreadable (an
// interrupted distribute never landed the sidecar) — both of which a fenced
// teardown preserves. The blanket path aborts rather than clobber a live
// newer mount, and the caller must guarantee no concurrent newer encode of
// this volume (e.g. the admin dedupe key, or an operator lock).
// - a non-zero value fences the teardown to strictly-older generations,
// preserving same-or-newer, generation 0, and an unreadable .vif — for a
// stale-worker cleanup that must never wipe a newer run's live shards.
//
// Returns ErrFullTeardownNotAcked (wrapped, so errors.Is matches) when a
// reachable server does not ack the full teardown.
//
// This is the single teardown primitive shared by the plugin-worker EC task
// and the shell ec.encode pre-cleanup, so the fence semantics cannot drift
// between the two paths.
func UnmountAndDeleteEcShards(
ctx context.Context,
dialOption grpc.DialOption,
server pb.ServerAddress,
collection string,
volumeID uint32,
shardIds []uint32,
encodeTsNs int64,
) error {
return operation.WithVolumeServerClient(false, server, dialOption,
func(client volume_server_pb.VolumeServerClient) error {
if _, err := client.VolumeEcShardsUnmount(ctx, &volume_server_pb.VolumeEcShardsUnmountRequest{
VolumeId: volumeID,
ShardIds: shardIds,
EncodeTsNs: encodeTsNs,
}); err != nil {
return fmt.Errorf("unmount: %w", err)
}
resp, err := client.VolumeEcShardsDelete(ctx, &volume_server_pb.VolumeEcShardsDeleteRequest{
VolumeId: volumeID,
Collection: collection,
ShardIds: shardIds,
FullTeardown: true,
EncodeTsNs: encodeTsNs,
})
if err != nil {
return fmt.Errorf("delete: %w", err)
}
if !resp.GetFullTeardownDone() {
return fmt.Errorf("delete on %s: %w", server, ErrFullTeardownNotAcked)
}
return nil
})
}
// EcFileGeneration parses the generation of a 2PC-staged <base>.v<N> file:
// -1 means the name is not a generation file of base.
func EcFileGeneration(name, base string) int64 {
suffix, ok := strings.CutPrefix(name, base+".v")
if !ok {
return -1
}
generation, err := strconv.ParseInt(suffix, 10, 64)
if err != nil || generation <= 0 {
return -1
}
return generation
}
// RemoveEcGenerationFiles removes 2PC generation files staged under base:
// <base>.ecNN.v<N>, <base>.ecx.v<N>, <base>.ecj.v<N>, <base>.ecsum.v<N> and
// <base>.vif.v<N>. generationsOlderThan == 0 removes every generation;
// otherwise only generations strictly below it. Returns the first real
// removal failure.
func RemoveEcGenerationFiles(baseFileName string, generationsOlderThan uint32) error {
var firstErr error
record := func(err error) {
if err != nil && firstErr == nil {
firstErr = err
}
}
dir, fileName := filepath.Dir(baseFileName), filepath.Base(baseFileName)
ecPrefix, vifName := fileName+".ec", fileName+".vif"
entries, err := os.ReadDir(dir)
if err != nil {
if os.IsNotExist(err) {
return nil
}
return err
}
for _, entry := range entries {
name := entry.Name()
// A generation file is <artifact>.v<N>; the last dot separates the
// staged-generation suffix from the artifact name.
artifact := name[:max(strings.LastIndexByte(name, '.'), 0)]
if artifact != vifName && !strings.HasPrefix(artifact, ecPrefix) {
continue
}
generation := EcFileGeneration(name, artifact)
if generation < 0 || (generationsOlderThan > 0 && generation >= int64(generationsOlderThan)) {
continue
}
if err := os.Remove(filepath.Join(dir, name)); err != nil && !os.IsNotExist(err) {
record(err)
}
}
return firstErr
}