Files
seaweedfs/weed/mount/peer_registrar_test.go
T
Chris Lu 8a6348d3e9 peer chunk sharing 4/8: mount registrar + HRW owner selection (#9133)
* proto: define MountRegister/MountList and MountPeer service

Adds the wire types for peer chunk sharing between weed mount clients:

* filer.proto: MountRegister / MountList RPCs so each mount can heartbeat
  its peer-serve address into a filer-hosted registry, and refresh the
  list of peers. Tiny payload; the filer stores only O(fleet_size) state.

* mount_peer.proto (new): ChunkAnnounce / ChunkLookup RPCs for the
  mount-to-mount chunk directory. Each fid's directory entry lives on
  an HRW-assigned mount; announces and lookups route to that mount.

No behavior yet — later PRs wire the RPCs into the filer and mount.
See design-weed-mount-peer-chunk-sharing.md for the full design.

* filer: add mount-server registry behind -peer.registry.enable

Implements tier 1 of the peer chunk sharing design: an in-memory registry
of live weed mount servers, keyed by peer address, refreshed by
MountRegister heartbeats and served by MountList.

* weed/filer/peer_registry.go: thread-safe map with TTL eviction; lazy
  sweep on List plus a background sweeper goroutine for bounded memory.

* weed/server/filer_grpc_server_peer.go: MountRegister / MountList RPC
  handlers. When -peer.registry.enable is false (the default), both RPCs
  are silent no-ops so probing older filers is harmless.

* -peer.registry.enable flag on weed filer; FilerOption.PeerRegistryEnabled
  wires it through.

Phase 1 is single-filer (no cross-filer replication of the registry);
mounts that fail over to another filer will re-register on the next
heartbeat, so the registry self-heals within one TTL cycle.

Part of the peer-chunk-sharing design; no behavior change at runtime
until a later PR enables the flag on both filer and mount.

* filer: nil-safe peerRegistryEnable + registry hardening

Addresses review feedback on PR #9131.

* Fix: nil pointer deref in the mini cluster. FilerOptions instances
  constructed outside weed/command/filer.go (e.g. miniFilerOptions in
  mini.go) do not populate peerRegistryEnable, so dereferencing the
  pointer panics at Filer startup. Use the same
  `nil && deref` idiom already used for distributedLock / writebackCache.

* Hardening (gemini review): registry now enforces three invariants:
  - empty peer_addr is silently rejected (no client-controlled sentinel
    mass-inserts)
  - TTL is capped at 1 hour so a runaway client cannot pin entries
  - new-entry count is capped at 10000 to bound memory; renewals of
    existing entries are always honored, so a full registry still
    heartbeats its existing members correctly

Covered by new unit tests.

* filer: rename -peer.registry.enable flag to -mount.p2p

Per review feedback: the old name "peer.registry.enable" leaked
the implementation ("registry") into the CLI surface. "mount.p2p"
is shorter and describes what it actually controls — whether this
filer participates in mount-to-mount peer chunk sharing.

Flag renames (all three keep default=true, idle cost is near-zero):
  -peer.registry.enable        ->  -mount.p2p         (weed filer)
  -filer.peer.registry.enable  ->  -filer.mount.p2p   (weed mini, weed server)

Internal variable names (mountPeerRegistryEnable, MountPeerRegistry)
keep their longer form — they describe the component, not the knob.

* filer: MountList returns DataCenter + List uses RLock

Two review follow-ups on the mount peer registry:

* weed/server/filer_grpc_server_mount_peer.go: MountList was dropping
  the DataCenter on the wire. The whole point of carrying DC separately
  from Rack is letting the mount-side fetcher re-rank peers by the
  two-level locality hierarchy (same-rack > same-DC > cross-DC); without
  DC in the response every remote peer collapsed to "unknown locality."

* weed/filer/mount_peer_registry.go: List() was taking a write lock so
  it could lazy-delete expired entries inline. But MountList is a
  read-heavy RPC hit on every mount's 30 s refresh loop, and Sweep is
  already wired as the sole reclamation path (same pattern as the
  mount-side PeerDirectory). Switch List to RLock + filter, let Sweep
  do the map mutation, so concurrent MountList callers don't serialize
  on each other.

Test updated to reflect the new contract (List no longer mutates the
map; Sweep is what drops expired entries).

* mount: add peer chunk sharing options + advertise address resolver

First cut at the peer chunk sharing wiring on the mount side. No
functional behavior yet — this PR just introduces the option fields,
the -peer.* flags, and the helper that resolves a reachable
host:port from them. The server implementation arrives in PR #5
(gRPC service) and the fetcher in PR #7.

* ResolvePeerAdvertiseAddr: an explicit -peer.advertise wins; else we
  use -peer.listen's bind host if specific; else util.DetectedHostAddress
  combined with the port. This is what gets registered with the filer
  and announced to peers, so wildcard binds no longer result in
  unreachable identities like "[::]:18080".

* Option fields: PeerEnabled, PeerListen, PeerAdvertise, PeerRack.
  One port handles both directory RPCs and streaming chunk fetches
  (see PR #1 FetchChunk proto), so there is no second -peer.grpc.*
  flag — the old HTTP byte-transfer path is gone.

* New flags on weed mount: -peer.enable, -peer.listen (default :18080),
  -peer.advertise (default auto), -peer.rack.

* mount: register with filer and maintain HRW seed view

Adds the mount-side tier-1 client. On startup the mount calls
MountRegister with its advertise address (PR #3) and keeps both the
filer entry and the local seed view fresh via background tickers
(30 s register / 30 s list, 90 s filer TTL).

* peer_hrw.go: pure rendezvous-hashing helper picking a single owner
  per fid via top-1 HRW. Adding or removing one seed moves only
  ~1/N fids.

* peer_registrar.go: heartbeat + list poller. Seeds() returns the
  slice directly (no per-call copy) since listOnce atomically swaps;
  background RPCs bind their context to Stop() so unmount doesn't
  hang on a slow filer.

* WFS wiring uses ResolvePeerAdvertiseAddr from PR #3 for the
  identity registered with the filer. No HTTP server, no second
  port — one reachable address represents the mount.

* mount: broadcast MountRegister/MountList to every filer

Previously the registrar called through wfs.WithFilerClient, which only
reaches whichever filer the WFS filer-client session happens to be on.
That meant two mounts pointing at different filers would never see each
other: the filer mount registries are in-memory and per-filer (no
filer-to-filer sync), so each mount's MountList only returned peers
that had also registered through the same filer.

This commit makes the registrar multi-filer aware:

  * NewPeerRegistrar now takes the full FilerAddresses slice and a
    per-filer dial function. The old single-filer peerFilerClient
    interface is gone.

  * registerOnce fans a MountRegister RPC out to every filer in
    parallel. Succeeds if at least one filer accepted — an unreachable
    filer is tolerated, logged, and retried on the next heartbeat.

  * listOnce polls every filer's MountList in parallel and merges the
    responses by peer_addr, keeping the newest LastSeenNs on duplicates.
    Mounts talking to different filers therefore converge once every
    filer has been polled once.

The merged-list property is what lets a fleet of mounts spread across
multiple filers still form a single HRW seed view. Each filer only ever
sees the subset of mounts that heartbeat through it, but the registrar
reconstructs the union client-side.

New unit tests guard both properties:
  - RegisterBroadcastsToAllFilers: one registerOnce hits all N filers.
  - ListMergesAcrossFilers: mount-a on filer-1 and mount-b on filer-2
    both appear in the merged seed set.
  - ListMergeKeepsNewestLastSeen: the same mount reported by two
    filers collapses to one entry with the freshest timestamp.
2026-04-18 20:03:45 -07:00

211 lines
7.0 KiB
Go

package mount
import (
"context"
"sync"
"testing"
"github.com/seaweedfs/seaweedfs/weed/pb"
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"google.golang.org/grpc"
)
// fakeFilerClient captures MountRegister/MountList calls and lets the test
// pre-seed MountList responses. Implements just enough of
// filer_pb.SeaweedFilerClient to drive the registrar.
type fakeFilerClient struct {
filer_pb.SeaweedFilerClient // embed for methods we don't need
mu sync.Mutex
registerCalls []filer_pb.MountRegisterRequest
listResponse filer_pb.MountListResponse
}
func (f *fakeFilerClient) MountRegister(ctx context.Context, req *filer_pb.MountRegisterRequest, opts ...grpc.CallOption) (*filer_pb.MountRegisterResponse, error) {
f.mu.Lock()
defer f.mu.Unlock()
f.registerCalls = append(f.registerCalls, *req)
return &filer_pb.MountRegisterResponse{}, nil
}
func (f *fakeFilerClient) MountList(ctx context.Context, req *filer_pb.MountListRequest, opts ...grpc.CallOption) (*filer_pb.MountListResponse, error) {
f.mu.Lock()
defer f.mu.Unlock()
resp := f.listResponse // value copy
return &resp, nil
}
// fakeFilerFleet maps each addr to its own fakeFilerClient so a test can
// simulate several filers with different registered/listed state.
type fakeFilerFleet struct {
clients map[pb.ServerAddress]*fakeFilerClient
}
func (f *fakeFilerFleet) dial(ctx context.Context, addr pb.ServerAddress, fn func(client filer_pb.SeaweedFilerClient) error) error {
c, ok := f.clients[addr]
if !ok {
// Treat unknown filer as "reachable but empty" — lets tests omit
// pre-populating a client when they don't care.
c = &fakeFilerClient{}
f.clients[addr] = c
}
return fn(c)
}
func singleFilerFleet(c *fakeFilerClient) ([]pb.ServerAddress, filerDialFn) {
addr := pb.ServerAddress("filer-1:18888")
fleet := &fakeFilerFleet{clients: map[pb.ServerAddress]*fakeFilerClient{addr: c}}
return []pb.ServerAddress{addr}, fleet.dial
}
func TestPeerRegistrar_StartPopulatesSeedsFromFiler(t *testing.T) {
fc := &fakeFilerClient{
listResponse: filer_pb.MountListResponse{
Mounts: []*filer_pb.MountInfo{
{PeerAddr: "mount-a:18080", Rack: "r1"},
{PeerAddr: "mount-b:18080", Rack: "r2"},
},
},
}
filers, dial := singleFilerFleet(fc)
r := NewPeerRegistrar(filers, dial, "self:18080", "dc1", "r1")
if err := r.registerOnce(context.Background()); err != nil {
t.Fatalf("registerOnce: %v", err)
}
if err := r.listOnce(context.Background()); err != nil {
t.Fatalf("listOnce: %v", err)
}
fc.mu.Lock()
if len(fc.registerCalls) != 1 {
t.Errorf("expected 1 register call, got %d", len(fc.registerCalls))
} else if fc.registerCalls[0].PeerAddr != "self:18080" {
t.Errorf("register sent wrong peer addr: %q", fc.registerCalls[0].PeerAddr)
}
fc.mu.Unlock()
seeds := r.Seeds()
if len(seeds) != 2 {
t.Errorf("expected 2 seeds, got %d", len(seeds))
}
owner := r.OwnerFor("3,01637037d6")
if owner != "mount-a:18080" && owner != "mount-b:18080" {
t.Errorf("OwnerFor returned unexpected addr: %q", owner)
}
}
func TestPeerRegistrar_HeartbeatTTLMatchesConfig(t *testing.T) {
fc := &fakeFilerClient{}
filers, dial := singleFilerFleet(fc)
r := NewPeerRegistrar(filers, dial, "self:18080", "", "")
if err := r.registerOnce(context.Background()); err != nil {
t.Fatalf("registerOnce: %v", err)
}
fc.mu.Lock()
defer fc.mu.Unlock()
if len(fc.registerCalls) != 1 {
t.Fatalf("expected 1 register call, got %d", len(fc.registerCalls))
}
want := int32(r.registerTTL.Seconds())
if got := fc.registerCalls[0].TtlSeconds; got != want {
t.Errorf("TtlSeconds: got %d want %d", got, want)
}
}
func TestPeerRegistrar_StopIsIdempotent(t *testing.T) {
filers, dial := singleFilerFleet(&fakeFilerClient{})
r := NewPeerRegistrar(filers, dial, "self:18080", "", "")
r.Stop()
r.Stop() // second call must be a no-op (no panic)
}
// TestPeerRegistrar_RegisterBroadcastsToAllFilers guards the core
// multi-filer property: a single registerOnce must hit every configured
// filer so mounts pointing at different filers still converge.
func TestPeerRegistrar_RegisterBroadcastsToAllFilers(t *testing.T) {
fc1 := &fakeFilerClient{}
fc2 := &fakeFilerClient{}
fc3 := &fakeFilerClient{}
a1, a2, a3 := pb.ServerAddress("f1:18888"), pb.ServerAddress("f2:18888"), pb.ServerAddress("f3:18888")
fleet := &fakeFilerFleet{clients: map[pb.ServerAddress]*fakeFilerClient{a1: fc1, a2: fc2, a3: fc3}}
r := NewPeerRegistrar([]pb.ServerAddress{a1, a2, a3}, fleet.dial, "self:18080", "", "")
if err := r.registerOnce(context.Background()); err != nil {
t.Fatalf("registerOnce: %v", err)
}
for addr, fc := range fleet.clients {
fc.mu.Lock()
if len(fc.registerCalls) != 1 {
t.Errorf("filer %s: got %d register calls, want 1", addr, len(fc.registerCalls))
}
fc.mu.Unlock()
}
}
// TestPeerRegistrar_ListMergesAcrossFilers guards cross-filer convergence:
// mount A registered on filer-1, mount B on filer-2; a registrar that
// lists both filers must see both mounts.
func TestPeerRegistrar_ListMergesAcrossFilers(t *testing.T) {
fc1 := &fakeFilerClient{
listResponse: filer_pb.MountListResponse{
Mounts: []*filer_pb.MountInfo{{PeerAddr: "mount-a:18080", Rack: "r1", LastSeenNs: 200}},
},
}
fc2 := &fakeFilerClient{
listResponse: filer_pb.MountListResponse{
Mounts: []*filer_pb.MountInfo{{PeerAddr: "mount-b:18080", Rack: "r2", LastSeenNs: 200}},
},
}
a1, a2 := pb.ServerAddress("f1:18888"), pb.ServerAddress("f2:18888")
fleet := &fakeFilerFleet{clients: map[pb.ServerAddress]*fakeFilerClient{a1: fc1, a2: fc2}}
r := NewPeerRegistrar([]pb.ServerAddress{a1, a2}, fleet.dial, "self:18080", "", "")
if err := r.listOnce(context.Background()); err != nil {
t.Fatalf("listOnce: %v", err)
}
seeds := r.Seeds()
if len(seeds) != 2 {
t.Fatalf("want 2 merged seeds, got %d: %+v", len(seeds), seeds)
}
addrs := map[string]bool{}
for _, s := range seeds {
addrs[s.PeerAddr] = true
}
if !addrs["mount-a:18080"] || !addrs["mount-b:18080"] {
t.Errorf("merged seeds missing an entry: %+v", addrs)
}
}
// TestPeerRegistrar_ListMergeKeepsNewestLastSeen guards the dedupe rule:
// the same mount reported by two filers collapses to one entry, keeping
// the freshest LastSeenNs for liveness-ordering decisions.
func TestPeerRegistrar_ListMergeKeepsNewestLastSeen(t *testing.T) {
fc1 := &fakeFilerClient{
listResponse: filer_pb.MountListResponse{
Mounts: []*filer_pb.MountInfo{{PeerAddr: "mount-a:18080", Rack: "r1", LastSeenNs: 100}},
},
}
fc2 := &fakeFilerClient{
listResponse: filer_pb.MountListResponse{
Mounts: []*filer_pb.MountInfo{{PeerAddr: "mount-a:18080", Rack: "r1", LastSeenNs: 500}},
},
}
a1, a2 := pb.ServerAddress("f1:18888"), pb.ServerAddress("f2:18888")
fleet := &fakeFilerFleet{clients: map[pb.ServerAddress]*fakeFilerClient{a1: fc1, a2: fc2}}
r := NewPeerRegistrar([]pb.ServerAddress{a1, a2}, fleet.dial, "self:18080", "", "")
if err := r.listOnce(context.Background()); err != nil {
t.Fatalf("listOnce: %v", err)
}
seeds := r.Seeds()
if len(seeds) != 1 {
t.Fatalf("want 1 deduped seed, got %d", len(seeds))
}
}