mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-09-08 15:41:15 +02:00
* proto: define MountRegister/MountList and MountPeer service Adds the wire types for peer chunk sharing between weed mount clients: * filer.proto: MountRegister / MountList RPCs so each mount can heartbeat its peer-serve address into a filer-hosted registry, and refresh the list of peers. Tiny payload; the filer stores only O(fleet_size) state. * mount_peer.proto (new): ChunkAnnounce / ChunkLookup RPCs for the mount-to-mount chunk directory. Each fid's directory entry lives on an HRW-assigned mount; announces and lookups route to that mount. No behavior yet — later PRs wire the RPCs into the filer and mount. See design-weed-mount-peer-chunk-sharing.md for the full design. * filer: add mount-server registry behind -peer.registry.enable Implements tier 1 of the peer chunk sharing design: an in-memory registry of live weed mount servers, keyed by peer address, refreshed by MountRegister heartbeats and served by MountList. * weed/filer/peer_registry.go: thread-safe map with TTL eviction; lazy sweep on List plus a background sweeper goroutine for bounded memory. * weed/server/filer_grpc_server_peer.go: MountRegister / MountList RPC handlers. When -peer.registry.enable is false (the default), both RPCs are silent no-ops so probing older filers is harmless. * -peer.registry.enable flag on weed filer; FilerOption.PeerRegistryEnabled wires it through. Phase 1 is single-filer (no cross-filer replication of the registry); mounts that fail over to another filer will re-register on the next heartbeat, so the registry self-heals within one TTL cycle. Part of the peer-chunk-sharing design; no behavior change at runtime until a later PR enables the flag on both filer and mount. * filer: nil-safe peerRegistryEnable + registry hardening Addresses review feedback on PR #9131. * Fix: nil pointer deref in the mini cluster. FilerOptions instances constructed outside weed/command/filer.go (e.g. miniFilerOptions in mini.go) do not populate peerRegistryEnable, so dereferencing the pointer panics at Filer startup. Use the same `nil && deref` idiom already used for distributedLock / writebackCache. * Hardening (gemini review): registry now enforces three invariants: - empty peer_addr is silently rejected (no client-controlled sentinel mass-inserts) - TTL is capped at 1 hour so a runaway client cannot pin entries - new-entry count is capped at 10000 to bound memory; renewals of existing entries are always honored, so a full registry still heartbeats its existing members correctly Covered by new unit tests. * filer: rename -peer.registry.enable flag to -mount.p2p Per review feedback: the old name "peer.registry.enable" leaked the implementation ("registry") into the CLI surface. "mount.p2p" is shorter and describes what it actually controls — whether this filer participates in mount-to-mount peer chunk sharing. Flag renames (all three keep default=true, idle cost is near-zero): -peer.registry.enable -> -mount.p2p (weed filer) -filer.peer.registry.enable -> -filer.mount.p2p (weed mini, weed server) Internal variable names (mountPeerRegistryEnable, MountPeerRegistry) keep their longer form — they describe the component, not the knob. * filer: MountList returns DataCenter + List uses RLock Two review follow-ups on the mount peer registry: * weed/server/filer_grpc_server_mount_peer.go: MountList was dropping the DataCenter on the wire. The whole point of carrying DC separately from Rack is letting the mount-side fetcher re-rank peers by the two-level locality hierarchy (same-rack > same-DC > cross-DC); without DC in the response every remote peer collapsed to "unknown locality." * weed/filer/mount_peer_registry.go: List() was taking a write lock so it could lazy-delete expired entries inline. But MountList is a read-heavy RPC hit on every mount's 30 s refresh loop, and Sweep is already wired as the sole reclamation path (same pattern as the mount-side PeerDirectory). Switch List to RLock + filter, let Sweep do the map mutation, so concurrent MountList callers don't serialize on each other. Test updated to reflect the new contract (List no longer mutates the map; Sweep is what drops expired entries). * mount: add peer chunk sharing options + advertise address resolver First cut at the peer chunk sharing wiring on the mount side. No functional behavior yet — this PR just introduces the option fields, the -peer.* flags, and the helper that resolves a reachable host:port from them. The server implementation arrives in PR #5 (gRPC service) and the fetcher in PR #7. * ResolvePeerAdvertiseAddr: an explicit -peer.advertise wins; else we use -peer.listen's bind host if specific; else util.DetectedHostAddress combined with the port. This is what gets registered with the filer and announced to peers, so wildcard binds no longer result in unreachable identities like "[::]:18080". * Option fields: PeerEnabled, PeerListen, PeerAdvertise, PeerRack. One port handles both directory RPCs and streaming chunk fetches (see PR #1 FetchChunk proto), so there is no second -peer.grpc.* flag — the old HTTP byte-transfer path is gone. * New flags on weed mount: -peer.enable, -peer.listen (default :18080), -peer.advertise (default auto), -peer.rack. * mount: register with filer and maintain HRW seed view Adds the mount-side tier-1 client. On startup the mount calls MountRegister with its advertise address (PR #3) and keeps both the filer entry and the local seed view fresh via background tickers (30 s register / 30 s list, 90 s filer TTL). * peer_hrw.go: pure rendezvous-hashing helper picking a single owner per fid via top-1 HRW. Adding or removing one seed moves only ~1/N fids. * peer_registrar.go: heartbeat + list poller. Seeds() returns the slice directly (no per-call copy) since listOnce atomically swaps; background RPCs bind their context to Stop() so unmount doesn't hang on a slow filer. * WFS wiring uses ResolvePeerAdvertiseAddr from PR #3 for the identity registered with the filer. No HTTP server, no second port — one reachable address represents the mount. * mount: broadcast MountRegister/MountList to every filer Previously the registrar called through wfs.WithFilerClient, which only reaches whichever filer the WFS filer-client session happens to be on. That meant two mounts pointing at different filers would never see each other: the filer mount registries are in-memory and per-filer (no filer-to-filer sync), so each mount's MountList only returned peers that had also registered through the same filer. This commit makes the registrar multi-filer aware: * NewPeerRegistrar now takes the full FilerAddresses slice and a per-filer dial function. The old single-filer peerFilerClient interface is gone. * registerOnce fans a MountRegister RPC out to every filer in parallel. Succeeds if at least one filer accepted — an unreachable filer is tolerated, logged, and retried on the next heartbeat. * listOnce polls every filer's MountList in parallel and merges the responses by peer_addr, keeping the newest LastSeenNs on duplicates. Mounts talking to different filers therefore converge once every filer has been polled once. The merged-list property is what lets a fleet of mounts spread across multiple filers still form a single HRW seed view. Each filer only ever sees the subset of mounts that heartbeat through it, but the registrar reconstructs the union client-side. New unit tests guard both properties: - RegisterBroadcastsToAllFilers: one registerOnce hits all N filers. - ListMergesAcrossFilers: mount-a on filer-1 and mount-b on filer-2 both appear in the merged seed set. - ListMergeKeepsNewestLastSeen: the same mount reported by two filers collapses to one entry with the freshest timestamp.
91 lines
2.5 KiB
Go
91 lines
2.5 KiB
Go
package mount
|
|
|
|
import (
|
|
"fmt"
|
|
"testing"
|
|
)
|
|
|
|
func mkSeeds(addrs ...string) []SeedPeer {
|
|
out := make([]SeedPeer, len(addrs))
|
|
for i, a := range addrs {
|
|
out[i] = SeedPeer{PeerAddr: a}
|
|
}
|
|
return out
|
|
}
|
|
|
|
func TestOwnerFor_EmptySeeds(t *testing.T) {
|
|
if got := OwnerFor("3,01637037d6", nil); got != "" {
|
|
t.Errorf("expected empty owner for empty seed list, got %q", got)
|
|
}
|
|
}
|
|
|
|
func TestOwnerFor_Deterministic(t *testing.T) {
|
|
seeds := mkSeeds("a:1", "b:1", "c:1", "d:1")
|
|
fid := "3,01637037d6"
|
|
first := OwnerFor(fid, seeds)
|
|
for i := 0; i < 100; i++ {
|
|
if got := OwnerFor(fid, seeds); got != first {
|
|
t.Fatalf("non-deterministic: iter %d got %q first %q", i, got, first)
|
|
}
|
|
}
|
|
}
|
|
|
|
func TestOwnerFor_DistributesEvenly(t *testing.T) {
|
|
const N = 4
|
|
seeds := mkSeeds("a:1", "b:1", "c:1", "d:1")
|
|
counts := map[string]int{}
|
|
for i := 0; i < 10000; i++ {
|
|
fid := fmt.Sprintf("%d,%x", i, i*17)
|
|
counts[OwnerFor(fid, seeds)]++
|
|
}
|
|
// Each seed should get roughly 25% (±10% slack for a 10k sample).
|
|
for addr, c := range counts {
|
|
ratio := float64(c) / 10000.0
|
|
if ratio < 0.225 || ratio > 0.275 {
|
|
t.Errorf("HRW distribution skewed: %s got %.3f", addr, ratio)
|
|
}
|
|
}
|
|
if len(counts) != N {
|
|
t.Errorf("expected %d distinct owners, got %d", N, len(counts))
|
|
}
|
|
}
|
|
|
|
func TestOwnerFor_MinimalShuffleOnSeedChange(t *testing.T) {
|
|
// Adding one seed should move ~1/(N+1) fids to the new seed and leave
|
|
// the rest on their prior owners. Tolerance generous for a 10k sample.
|
|
const trials = 10000
|
|
before := mkSeeds("a:1", "b:1", "c:1")
|
|
after := mkSeeds("a:1", "b:1", "c:1", "d:1")
|
|
|
|
moved := 0
|
|
toNewSeed := 0
|
|
for i := 0; i < trials; i++ {
|
|
fid := fmt.Sprintf("%d,%x", i, i*31)
|
|
pre := OwnerFor(fid, before)
|
|
post := OwnerFor(fid, after)
|
|
if pre != post {
|
|
moved++
|
|
if post == "d:1" {
|
|
toNewSeed++
|
|
}
|
|
}
|
|
}
|
|
// Expected: ~1/4 = 25% of fids move, all to the new seed.
|
|
ratio := float64(moved) / float64(trials)
|
|
if ratio < 0.20 || ratio > 0.30 {
|
|
t.Errorf("expected ~25%% fids to move on seed-add, got %.3f", ratio)
|
|
}
|
|
if moved != toNewSeed {
|
|
t.Errorf("expected every moved fid to land on the new seed; moved=%d toNewSeed=%d", moved, toNewSeed)
|
|
}
|
|
}
|
|
|
|
func TestOwnerFor_TieBreakerDeterministic(t *testing.T) {
|
|
// Two seeds with equal hash score (engineered collision is unlikely in
|
|
// practice, but we guarantee determinism by lex-comparing addresses).
|
|
// This test just confirms OwnerFor runs without panicking on duplicates.
|
|
seeds := mkSeeds("a:1", "a:1", "b:1")
|
|
fid := "3,01637037d6"
|
|
_ = OwnerFor(fid, seeds)
|
|
}
|