mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-06 14:31:57 +02:00
* volume: remove staged EC generation files on teardown and shard delete The 2PC generation switch stages each run as <base>.ecNN.v<N> plus versioned .ecx/.ecj/.vif files. Nothing on the volume server removes them: isEcDataShardFile only recognises the exact .ecNN name, so the staged files are invisible to every bookkeeping pass, and even full_teardown's wipe-all path left them behind. Each re-encode therefore leaks a full shard set per shard-holding disk. RemoveEcGenerationFiles sweeps <base>.ec*.v<N> and <base>.vif.v<N>, optionally keeping generations at or above a threshold; teardown and the reconcile wipe remove every generation, and a per-shard delete removes that shard's staged generations too. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume: delete staged EC generations older than N via VolumeEcShardsDelete After a 2PC generation switch commits, the superseded generation's <base>.*.v<N> files sit on disk with no cleanup path: teardown removes everything, and a per-shard delete only touches the named shards, so the executor had no RPC that reclaims just the staged leftovers. delete_generations_older_than removes staged generation files strictly below the threshold on every disk. Versioned files are never mounted, so nothing is unloaded first; the committed generation and the canonical files are preserved. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: mirror staged EC generation cleanup Parity with the Go volume server: remove_ec_generation_files sweeps <base>.ec*.v<N> and <base>.vif.v<N> staged by the 2PC switch, called by remove_ec_volume_files (which covers both teardown paths) and the new delete_generations_older_than request field; delete_ec_shards removes a shard's staged generations along with the canonical file. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume: match staged generation filenames literally filepath.Glob interprets metacharacters in the collection part of the base name, so a collection like a[bc] could match another volume's staged files (or miss its own). Scan the directory and compare names literally instead, mirroring the Rust read_dir implementation. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: report generation-sweep errors and drop the store lock first - snapshot the location base names under the read lock and run the filesystem sweep after dropping it, so a slow disk cannot stall the store; - record per-entry read_dir errors in remove_ec_generation_files and propagate them from remove_ec_shard_generations instead of flatten() skipping them; - warn when a staged-shard generation fails to delete rather than reporting success with files left behind. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume: fail shard delete when the staged-generation listing fails A transient ReadDir failure fell back to removing canonical shard names only: staged .v<N> files survived while the RPC still reported success, leaving the leak invisible to retrying callers. ENOENT still means the disk simply has no such directory; other listing errors now propagate. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: propagate staged-generation removal failures delete_ec_shards logged remove_ec_shard_generations errors and the RPC returned success while staged .v<N> files remained, diverging from the Go handler which surfaces the failure. The sweep keeps processing the remaining shards, retains the first error, and volume_ec_shards_delete maps it to Status::internal so callers can retry. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: notify state change even when the shard sweep errors delete_ec_shards already deletes and unmounts the shards before returning a staged-generation failure, so returning early skipped volume_state_notify and the master kept routing to them until the next heartbeat. Notify before propagating the error. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
135 lines
4.8 KiB
Go
135 lines
4.8 KiB
Go
package erasure_coding
|
|
|
|
import (
|
|
"context"
|
|
"errors"
|
|
"fmt"
|
|
"os"
|
|
"path/filepath"
|
|
"strconv"
|
|
"strings"
|
|
|
|
"github.com/seaweedfs/seaweedfs/weed/operation"
|
|
"github.com/seaweedfs/seaweedfs/weed/pb"
|
|
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
|
|
"google.golang.org/grpc"
|
|
)
|
|
|
|
// ErrFullTeardownNotAcked marks a reachable server that completed the delete
|
|
// RPC but did not report a full teardown (e.g. a pre-upgrade volume server), so
|
|
// a stale EC generation may remain on it. Callers distinguish this from an
|
|
// unreachable node (which may recover and be re-swept) with errors.Is.
|
|
var ErrFullTeardownNotAcked = errors.New("delete did not perform full teardown (pre-upgrade volume server?); a stale EC generation may remain")
|
|
|
|
// UnmountAndDeleteEcShards unmounts then tears down the named EC shards for a
|
|
// volume on one server. Unmount must precede delete (delete requires the shard
|
|
// be unmounted); both RPCs are idempotent against missing shards.
|
|
//
|
|
// encodeTsNs fences both RPCs:
|
|
// - 0 selects the server's blanket, generation-independent teardown. This is
|
|
// the correct choice for a pre-encode or rollback wipe: it clears same-
|
|
// generation shards (a retried encode's prior attempt shares the job's
|
|
// generation) and shards whose .vif generation is unreadable (an
|
|
// interrupted distribute never landed the sidecar) — both of which a fenced
|
|
// teardown preserves. The blanket path aborts rather than clobber a live
|
|
// newer mount, and the caller must guarantee no concurrent newer encode of
|
|
// this volume (e.g. the admin dedupe key, or an operator lock).
|
|
// - a non-zero value fences the teardown to strictly-older generations,
|
|
// preserving same-or-newer, generation 0, and an unreadable .vif — for a
|
|
// stale-worker cleanup that must never wipe a newer run's live shards.
|
|
//
|
|
// Returns ErrFullTeardownNotAcked (wrapped, so errors.Is matches) when a
|
|
// reachable server does not ack the full teardown.
|
|
//
|
|
// This is the single teardown primitive shared by the plugin-worker EC task
|
|
// and the shell ec.encode pre-cleanup, so the fence semantics cannot drift
|
|
// between the two paths.
|
|
func UnmountAndDeleteEcShards(
|
|
ctx context.Context,
|
|
dialOption grpc.DialOption,
|
|
server pb.ServerAddress,
|
|
collection string,
|
|
volumeID uint32,
|
|
shardIds []uint32,
|
|
encodeTsNs int64,
|
|
) error {
|
|
return operation.WithVolumeServerClient(false, server, dialOption,
|
|
func(client volume_server_pb.VolumeServerClient) error {
|
|
if _, err := client.VolumeEcShardsUnmount(ctx, &volume_server_pb.VolumeEcShardsUnmountRequest{
|
|
VolumeId: volumeID,
|
|
ShardIds: shardIds,
|
|
EncodeTsNs: encodeTsNs,
|
|
}); err != nil {
|
|
return fmt.Errorf("unmount: %w", err)
|
|
}
|
|
resp, err := client.VolumeEcShardsDelete(ctx, &volume_server_pb.VolumeEcShardsDeleteRequest{
|
|
VolumeId: volumeID,
|
|
Collection: collection,
|
|
ShardIds: shardIds,
|
|
FullTeardown: true,
|
|
EncodeTsNs: encodeTsNs,
|
|
})
|
|
if err != nil {
|
|
return fmt.Errorf("delete: %w", err)
|
|
}
|
|
if !resp.GetFullTeardownDone() {
|
|
return fmt.Errorf("delete on %s: %w", server, ErrFullTeardownNotAcked)
|
|
}
|
|
return nil
|
|
})
|
|
}
|
|
|
|
// EcFileGeneration parses the generation of a 2PC-staged <base>.v<N> file:
|
|
// -1 means the name is not a generation file of base.
|
|
func EcFileGeneration(name, base string) int64 {
|
|
suffix, ok := strings.CutPrefix(name, base+".v")
|
|
if !ok {
|
|
return -1
|
|
}
|
|
generation, err := strconv.ParseInt(suffix, 10, 64)
|
|
if err != nil || generation <= 0 {
|
|
return -1
|
|
}
|
|
return generation
|
|
}
|
|
|
|
// RemoveEcGenerationFiles removes 2PC generation files staged under base:
|
|
// <base>.ecNN.v<N>, <base>.ecx.v<N>, <base>.ecj.v<N>, <base>.ecsum.v<N> and
|
|
// <base>.vif.v<N>. generationsOlderThan == 0 removes every generation;
|
|
// otherwise only generations strictly below it. Returns the first real
|
|
// removal failure.
|
|
func RemoveEcGenerationFiles(baseFileName string, generationsOlderThan uint32) error {
|
|
var firstErr error
|
|
record := func(err error) {
|
|
if err != nil && firstErr == nil {
|
|
firstErr = err
|
|
}
|
|
}
|
|
dir, fileName := filepath.Dir(baseFileName), filepath.Base(baseFileName)
|
|
ecPrefix, vifName := fileName+".ec", fileName+".vif"
|
|
entries, err := os.ReadDir(dir)
|
|
if err != nil {
|
|
if os.IsNotExist(err) {
|
|
return nil
|
|
}
|
|
return err
|
|
}
|
|
for _, entry := range entries {
|
|
name := entry.Name()
|
|
// A generation file is <artifact>.v<N>; the last dot separates the
|
|
// staged-generation suffix from the artifact name.
|
|
artifact := name[:max(strings.LastIndexByte(name, '.'), 0)]
|
|
if artifact != vifName && !strings.HasPrefix(artifact, ecPrefix) {
|
|
continue
|
|
}
|
|
generation := EcFileGeneration(name, artifact)
|
|
if generation < 0 || (generationsOlderThan > 0 && generation >= int64(generationsOlderThan)) {
|
|
continue
|
|
}
|
|
if err := os.Remove(filepath.Join(dir, name)); err != nil && !os.IsNotExist(err) {
|
|
record(err)
|
|
}
|
|
}
|
|
return firstErr
|
|
}
|