mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-08 07:17:48 +02:00
p2p
1 parent
3116bdeec6
commit
674f0efdfb
2 files changed
+163
No files matched your search
@@ -0,0 +1,162 @@
|
||||
# Peer Chunk Sharing Between weed mount Clients
|
||||
|
||||
When a fleet of GPU hosts loads the same model file through `weed mount`, every client pulls bytes from the volume tier. Even with [`fs.distributeChunks`](Distributing-AI-Model-Files-for-Multi-GPU-Loading) spreading the chunks evenly across every volume server, the volume tier's total NIC bandwidth caps how fast the read burst can complete. Peer chunk sharing lets mounts fetch chunks from each other instead, so after the first wave seeds a handful of mounts, subsequent reads fan out across the whole fleet.
|
||||
|
||||
This feature is opt-in via `-peer.enable=true` on `weed mount`. On the filer side it is enabled by default via `-mount.p2p=true` (idle cost is negligible). When the mount flag is off — the default for `weed mount` — reads behave exactly as they do today.
|
||||
|
||||
The design is documented in [design-weed-mount-peer-chunk-sharing.md](https://github.com/seaweedfs/seaweedfs/blob/master/design-weed-mount-peer-chunk-sharing.md). This page is the operator-facing summary.
|
||||
|
||||
## How it works at a glance
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Filer["Filer(s)<br/>mount registry"]
|
||||
MountA["weed mount A"]
|
||||
MountB["weed mount B"]
|
||||
MountC["weed mount C"]
|
||||
Volumes["Volume servers"]
|
||||
|
||||
MountA -- "MountRegister / MountList" --> Filer
|
||||
MountB -- "MountRegister / MountList" --> Filer
|
||||
MountC -- "MountRegister / MountList" --> Filer
|
||||
|
||||
MountA <-- "ChunkAnnounce / Lookup / FetchChunk" --> MountB
|
||||
MountB <-- "ChunkAnnounce / Lookup / FetchChunk" --> MountC
|
||||
MountA <-- "ChunkAnnounce / Lookup / FetchChunk" --> MountC
|
||||
|
||||
MountA -. "fallback on any failure" .-> Volumes
|
||||
MountB -. "fallback on any failure" .-> Volumes
|
||||
MountC -. "fallback on any failure" .-> Volumes
|
||||
```
|
||||
|
||||
* **Tier 1 — filer registry.** Each filer holds a tiny in-memory map of which mounts are alive. Mounts broadcast `MountRegister` to **every** configured filer, and merge every filer's `MountList` response by `peer_addr` (newest `last_seen_ns` wins). That way two mounts pointing at different filers still see each other even though no filer-to-filer sync exists.
|
||||
* **Tier 2 — mount-hosted chunk directory.** The mount fleet itself shards a `fid → holders` directory via rendezvous hashing (HRW) on the registered mount list. Each owner mount only holds directory entries for fids it's HRW-assigned.
|
||||
* **One gRPC port.** `ChunkAnnounce`, `ChunkLookup`, and the chunk-byte stream (`FetchChunk`) all go through a single mount-to-mount gRPC service. No separate HTTP peer-serve port.
|
||||
* **Peer-serve path.** On the read path, each mount asks the HRW owner for holders, picks the best peer by locality (same rack > same DC > elsewhere, LRU tiebreak within each bucket), and server-streams the chunk bytes back. MD5/ETag verification on the full assembled buffer before returning to the kernel; any failure falls through cleanly to the volume tier.
|
||||
|
||||
## Enabling the feature
|
||||
|
||||
### On the filer
|
||||
|
||||
The filer accepts `MountRegister` / `MountList` RPCs by default. The knob is `-mount.p2p` (default `true`); set it to `false` if you want to opt out, or if you're running an older mount fleet that shouldn't be discoverable yet.
|
||||
|
||||
```bash
|
||||
weed filer -mount.p2p=true \
|
||||
-master=master1:9333,master2:9333,master3:9333 \
|
||||
-ip=filer1 -port=8888
|
||||
```
|
||||
|
||||
The same flag is available as `-filer.mount.p2p` on `weed mini` and `weed server`.
|
||||
|
||||
### On each mount
|
||||
|
||||
```bash
|
||||
weed mount -filer=filer1:8888,filer2:8888 \
|
||||
-dir=/mnt/seaweedfs \
|
||||
-peer.enable=true \
|
||||
-peer.listen=:18080 \
|
||||
-peer.advertise=10.0.0.5:18080 \
|
||||
-peer.dataCenter=dc-east \
|
||||
-peer.rack=rack-a
|
||||
```
|
||||
|
||||
The mount will:
|
||||
|
||||
1. Register with **every** configured filer's mount registry and heartbeat every 30 s.
|
||||
2. Listen on `:18080` (gRPC) for `ChunkAnnounce` / `ChunkLookup` / `FetchChunk`.
|
||||
3. On reads, try peer mounts before the volume tier. On any failure (owner unreachable, holder has since evicted, ETag mismatch, etc.) it falls through transparently to the volume path.
|
||||
|
||||
`-peer.advertise` is optional; when set, mounts receive that address from the filer's `MountList` instead of whatever the bind string would resolve to. It is required when `-peer.listen` uses a wildcard host (`":18080"`, `"0.0.0.0:18080"`, `"[::]:18080"`) and auto-detection can't find a reachable IP — in that case the mount will fail to start rather than advertise an unusable loopback address.
|
||||
|
||||
### Turning it off
|
||||
|
||||
Set `-peer.enable=false` (or omit the flag) on a rolling restart. The feature disables cleanly — the mount stops serving peers, stops registering with filers, and reads take the same path they did before.
|
||||
|
||||
## Flag reference
|
||||
|
||||
### `weed mount`
|
||||
|
||||
| Flag | Default | Meaning |
|
||||
|------|---------|---------|
|
||||
| `-peer.enable` | `false` | Opt-in master switch. |
|
||||
| `-peer.listen` | `:18080` | bind address for the peer gRPC server (directory RPCs + chunk streaming). |
|
||||
| `-peer.advertise` | *(auto-detect)* | externally-reachable `host:port` other mounts use to reach this one. Required with a wildcard `-peer.listen` when auto-detect fails. |
|
||||
| `-peer.dataCenter` | `""` | data-center locality label advertised to peers. |
|
||||
| `-peer.rack` | `""` | rack locality label (finer than DC). |
|
||||
|
||||
### `weed filer`
|
||||
|
||||
| Flag | Default | Meaning |
|
||||
|------|---------|---------|
|
||||
| `-mount.p2p` | `true` | Accept `MountRegister` / `MountList` from mount clients. |
|
||||
|
||||
Same flag on `weed mini` and `weed server` is namespaced as `-filer.mount.p2p`.
|
||||
|
||||
## When it helps
|
||||
|
||||
- Many mount clients reading the same large file in overlapping windows — LLM model loading across a GPU fleet is the canonical case.
|
||||
- Fleets where inter-host bandwidth (10/25/100 GbE between GPU hosts) is abundant but the volume-tier aggregate NIC bandwidth is the bottleneck.
|
||||
- Workloads with long-lived content access: chunks that stay cached on one mount remain servable to new arrivals for as long as they sit in that mount's on-disk cache.
|
||||
|
||||
## When it does not help
|
||||
|
||||
- Single-mount workloads: no peers to share with.
|
||||
- Short-lived mount processes whose caches evict quickly: the TTL window for being discoverable is short.
|
||||
- Writes: chunks are not shared peer-to-peer during writes; writes always go to volume servers. Peer sharing is read-only.
|
||||
|
||||
## Operational notes
|
||||
|
||||
### Port and firewall
|
||||
|
||||
Each mount binds `-peer.listen` for its gRPC server. That port must be reachable by every other mount in the same cluster (it carries directory RPCs and the chunk-byte stream). On Kubernetes, typically a `ClusterIP` service or `hostNetwork` pods; on bare metal, open the port on the inter-host network only.
|
||||
|
||||
### Authentication
|
||||
|
||||
The peer gRPC service reuses the same transport credentials the mount already uses for talking to the filer. When `security.toml` configures gRPC TLS, peer connections use it too — no separate credential.
|
||||
|
||||
### Integrity verification
|
||||
|
||||
Every fetched peer response is verified end-to-end by MD5 against `FileChunk.ETag` from the filer entry before its bytes are handed to the kernel. Mismatch → discard, fall through to the volume tier. This closes the trust gap opened by treating peer mounts as untrusted sources.
|
||||
|
||||
### Cache and TTL
|
||||
|
||||
Directory entries on owner mounts expire after 300 s (5 min) without a renewing `ChunkAnnounce`. Holder mounts re-announce fids they still hold roughly once per 270 s. The TTL is tuned for the desynchronized loader pattern where chunks stay cached for hours; bursty fleets that want faster eviction can shorten the interval in a later configuration flag.
|
||||
|
||||
No explicit retraction is sent on eviction — stale directory entries return a gRPC `NOT_FOUND` from the peer's `FetchChunk` call and the caller falls through. A single wasted RTT, no correctness impact.
|
||||
|
||||
### Locality-aware peer selection
|
||||
|
||||
When the owner returns multiple holders, the fetcher re-ranks them client-side:
|
||||
|
||||
1. Same rack (same `-peer.rack` AND same `-peer.dataCenter`) — best.
|
||||
2. Same DC, different rack — next best.
|
||||
3. Cross-DC or unknown labels — last.
|
||||
|
||||
Within each bucket the server's LRU order is preserved (freshest holder first). Unlabeled peers end up in bucket 3 — always give every mount at least `-peer.dataCenter` if you want meaningful locality ranking.
|
||||
|
||||
### Multi-filer deployments
|
||||
|
||||
The registrar broadcasts `MountRegister` to every filer listed in `-filer=` and merges every filer's `MountList` response. That's what lets two mounts pointing at different filers find each other even though the filer registries are in-memory per-filer with no cross-filer sync. An unreachable filer is tolerated; the mount keeps running as long as at least one filer succeeded on the last heartbeat.
|
||||
|
||||
## Disabling in an emergency
|
||||
|
||||
If peer sharing misbehaves in production, the kill switch is a rolling restart with `-peer.enable=false` on the mounts. Because the read path falls through to the volume tier on every failure mode, you should not observe read errors during the restart — just a gradual transfer of load back to the volume servers.
|
||||
|
||||
If needed, the filer side can be disabled by setting `-mount.p2p=false`. Existing mounts' heartbeats become no-ops and the peer-fetch path's `ChunkLookup` calls return empty within seconds (still falling through cleanly).
|
||||
|
||||
## Limitations
|
||||
|
||||
- **No write-path announce**: mounts don't advertise chunks they just uploaded. Same-host write-then-read still works (local cache), but cross-mount discovery of freshly-written chunks waits until another mount reads them first.
|
||||
- **Chunk manifests**: supported transparently — the fetcher resolves manifests to leaf chunks before the HRW lookup, so large files using manifest indirection participate in peer sharing at the leaf level.
|
||||
- **No metrics port on `weed mount`**: internal counters exist but the mount command does not currently expose a Prometheus endpoint. Observability lives in the glog stream (V(2) for warnings, V(4) for per-read success/failure).
|
||||
|
||||
## End-to-end CI
|
||||
|
||||
`test/fuse_p2p/` contains a FUSE-backed integration test that brings up a cluster of 3 mounts with `-peer.enable`, writes a file through one, and verifies a second mount can satisfy the read from the first mount's chunk cache. The `.github/workflows/fuse-p2p-integration.yml` workflow runs it on every pull request touching mount or filer peer code.
|
||||
|
||||
## See also
|
||||
|
||||
- [Distributing AI Model Files for Multi-GPU Loading](Distributing-AI-Model-Files-for-Multi-GPU-Loading) — the placement side of the same problem; run `fs.distributeChunks` after upload before letting peer sharing do its thing.
|
||||
- [FUSE Mount](FUSE-Mount) — baseline mount configuration.
|
||||
- [weed shell](weed-shell) — operator console.
|
||||
- Full design: `design-weed-mount-peer-chunk-sharing.md` in the SeaweedFS repo root.
|
||||
+1
@@ -132,6 +132,7 @@
|
||||
### Machine Learning
|
||||
* [[TensorFlow with SeaweedFS]]
|
||||
* [[Distributing AI Model Files for Multi-GPU Loading]]
|
||||
* [[Peer Chunk Sharing Between weed mount Clients]]
|
||||
|
||||
### HDFS
|
||||
* [[Hadoop Compatible File System]]
|
||||
|
||||
Reference in new issue
Block a user