Files
seaweedfs/weed/operation/volume_move/mover.go
T
Chris Lu 0799084e98 refactor: share volume and EC shard move logic between shell and workers (#10727)
* operation: add shared volume_move package for volume and EC shard moves

The shell commands (volume.move, volume.balance, ec.balance, tier moves)
and the maintenance workers (balance, ec_balance) each carried their own
copy of the move RPC sequences, and the copies had drifted: the worker
verified the target before deleting the source but dropped the disk
type and IO throttle; the shell passed those but deleted the source
unverified.

volume_move.Mover carries the merged sequences, keeping the stricter
behavior from each side:

- LiveMoveVolume: check-then-hard-freeze the source (VolumeStatus's
  IsReadOnly also covers low-disk and readonly-but-can-delete states,
  which still accept needle deletes), copy with disk type and IO
  throttle, tail, verify the target is not behind the source before the
  destructive source delete (a target that is ahead holds writes it
  accepted during the tail and the move commits to keep them), and
  restore the source's writability when a failure precedes the delete
  and this move did the freezing. Aborts clean up the incomplete target
  copy; a failed cleanup or an ambiguous source delete keeps the source
  readonly (ErrSourceKeptReadonly) so callers do not thaw a source next
  to a possibly-authoritative copy. With a readonly source, an existing
  or unknown-state target refuses the move outright: no client-side
  observation can prove such a copy is a stale remnant rather than the
  authoritative copy of an unfinished move.
- MoveEcShards: copy with the .ecx/.ecj/.vif/.ecsum sidecars, mount,
  verify the target registered every shard before unmount+delete on the
  source, and reject same-server moves (the EC delete is server-wide).

Server identity is the grpc endpoint (SameServer), so node:8080 and
node:8080.18080 compare equal while test servers sharing a degenerate
HTTP address stay distinct; addresses are validated non-fatally before
dialing and before being embedded in copy/tail requests, since both the
client dialer and the receiving server normalize them through a parser
that aborts the process on a malformed port. The Rust volume server's
codes.NotFound counts as a definitively absent probe answer alongside
the Go server's plain-error code Unknown.

All RPCs go through an injectable ClientFunc, so the sequences are unit
tested against a fake volume server client: RPC order, request fields,
and that verification failures keep the source intact.

* shell, worker: delegate volume and EC shard moves to operation/volume_move

LiveMoveVolume and the copy/tail/delete/mark-writable helpers become
thin wrappers over the shared mover, keeping their signatures; the EC
helpers keep their per-step output and delegate the RPCs. BalanceTask
and ECBalanceTask keep their parameter validation, progress reporting,
and guards (same-node cross-disk rejection, dedup keep-node
verification, shard ids range-checked before the uint8 narrowing) and
hand the RPC sequences to the mover. volume.tier.move skips its
thaw-on-failure when the mover deliberately kept the source readonly,
since reopening the replicas beside a possibly-authoritative target
copy would fork the volume.

The tail-failure tolerance moves inside the mover: a failed tail is
tolerated only when the volume was already readonly before the move
began, backstopped by a stability re-read across the idle window, so
volume.balance's -skipTailError-by-readonly heuristic and tier-move's
unconditional skip both become the same authoritative rule.

* volume_move: keep the source readonly when a failed copy leaves a target of unknown origin

A failed copy can leave a complete, mounted copy on the target (the
server finishes after the client loses the stream). The abort probed
the target only when its pre-copy state was known-absent; an unknown
prior state skipped both the probe and the cleanup and then reopened
the source - two writable replicas of one volume, diverging from the
next write on.

The abort now probes the target on every failed copy and restores the
source only when the target provably holds nothing. A copy whose
provenance cannot be proven (unknown prior state, a pre-existing
replica, or an unreachable target) is never deleted, and the source
stays readonly with ErrSourceKeptReadonly naming the recovery.

* test: teach the plugin worker harness the shared move sequence

The fake volume server lacked VolumeStatus, which the shared mover now
issues before freezing the source, and the batch execution test's
status-read accounting predates the pre-copy target probe and the
verification reads. Mirrors the harness the enterprise tree already
carries.
2026-08-12 12:29:40 -07:00

114 lines
4.7 KiB
Go

// Package volume_move implements the volume and EC shard move sequences shared
// by the interactive shell commands (volume.move, volume.balance, ec.balance,
// volume.tier.move, ...) and the maintenance workers (balance, ec_balance,
// volume_tiering). The RPCs go through an injectable ClientFunc so every
// sequence can be unit tested against a fake volume server client.
package volume_move
import (
"fmt"
"strconv"
"strings"
"github.com/seaweedfs/seaweedfs/weed/operation"
"github.com/seaweedfs/seaweedfs/weed/pb"
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
"github.com/seaweedfs/seaweedfs/weed/util"
"google.golang.org/grpc"
)
// ClientFunc runs fn with a client for the volume server at addr. It matches
// operation.WithVolumeServerClient, which production movers use to dial real
// servers; tests substitute a fake client.
type ClientFunc func(streamingMode bool, addr pb.ServerAddress, fn func(client volume_server_pb.VolumeServerClient) error) error
// Mover executes volume and EC shard moves against volume servers.
type Mover struct {
withClient ClientFunc
}
func NewMover(grpcDialOption grpc.DialOption) *Mover {
return &Mover{withClient: func(streamingMode bool, addr pb.ServerAddress, fn func(client volume_server_pb.VolumeServerClient) error) error {
// Validated here, the single point every mover RPC dials through:
// handing a malformed address (an unvalidated -source/-target flag)
// to the dialer would abort the whole process instead of failing the
// move.
if err := checkDialable(addr); err != nil {
return err
}
return operation.WithVolumeServerClient(streamingMode, addr, grpcDialOption, fn)
}}
}
// checkDialable rejects an address whose grpc normalization would abort the
// process: for the "host:port" form, ServerAddress.ToGrpcAddress falls back
// to a parser that calls glog.Fatalf when the port is not numeric. Everything
// else either normalizes cleanly or fails at dial time as an ordinary error.
// It also guards the source addresses embedded in copy/tail requests: the
// receiving volume server dials those through the same fatal parser, so an
// unchecked malformed source would terminate the destination server.
func checkDialable(addr pb.ServerAddress) error {
s := string(addr)
colon := strings.LastIndex(s, ":")
if colon < 0 || colon+1 >= len(s) {
return nil // no port part; handed to the dialer untouched
}
ports := s[colon+1:]
if dot := strings.LastIndex(ports, "."); dot >= 0 {
// "port.grpcPort": different dial paths canonicalize a half-malformed
// form differently (the method splices the grpc part in unparsed; the
// string helper falls back to the http part + 10000), so a bad
// component could dial an unintended server or reach the fatal
// parser. Require both components numeric.
if _, err := strconv.ParseUint(ports[:dot], 10, 64); err != nil {
return fmt.Errorf("invalid volume server address %q: port %q is not a number", s, ports[:dot])
}
if _, err := strconv.ParseUint(ports[dot+1:], 10, 64); err != nil {
return fmt.Errorf("invalid volume server address %q: grpc port %q is not a number", s, ports[dot+1:])
}
return nil
}
if _, err := strconv.ParseUint(ports, 10, 64); err != nil {
return fmt.Errorf("invalid volume server address %q: port %q is not a number", s, ports)
}
return nil
}
// NewMoverWithClientFunc builds a Mover on a custom transport.
func NewMoverWithClientFunc(withClient ClientFunc) *Mover {
return &Mover{withClient: withClient}
}
// SameServer reports whether two addresses name the same volume server. The
// gRPC endpoint identifies the server process — each has exactly one — so
// "node:8080" and "node:8080.18080" compare equal, while two servers sharing
// a degenerate HTTP address (e.g. port 0 in test harnesses) stay distinct.
// The normalization is non-fatal, unlike ServerAddress.ToGrpcAddress, whose
// parser exits the process on a malformed port; anything unparsable compares
// as its literal self.
func SameServer(a, b pb.ServerAddress) bool {
return grpcEndpoint(string(a)) == grpcEndpoint(string(b))
}
// grpcEndpoint mirrors ServerAddress.ToGrpcAddress ("host:port" gets the
// +10000 grpc port; "host:port.grpcPort" names it explicitly) but returns a
// malformed address unchanged instead of exiting.
func grpcEndpoint(addr string) string {
colon := strings.LastIndex(addr, ":")
if colon < 0 || colon+1 >= len(addr) {
return addr
}
host, ports := addr[:colon], addr[colon+1:]
if dot := strings.LastIndex(ports, "."); dot >= 0 {
if grpcPort, err := strconv.Atoi(ports[dot+1:]); err == nil {
return util.JoinHostPort(host, grpcPort)
}
return addr
}
httpPort, err := strconv.Atoi(ports)
if err != nil {
return addr
}
return util.JoinHostPort(host, httpPort+10000)
}