mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-08 15:27:43 +02:00
Document new topology-aware EC rebalancing logic.
1 parent
1d84deb0a8
commit
dedc5d9648
1 file changed
+9
-1
@@ -47,7 +47,7 @@ The scripts have 3 steps related to erasure coding.
|
||||
### Erasure Encode Sealed Data
|
||||
`ec.encode` command will find volumes that are almost full and has been stale for a period of time.
|
||||
|
||||
The default command is `ec.encode -fullPercent=95 -quietFor=1h`. It will find volumes at least 95% of the maximum volume size, which is usually 30GB, and have no updates for 1 hour.
|
||||
The default command is `ec.encode -fullPercent=95 -quietFor=1h -rebalance`. It will find volumes at least 95% of the maximum volume size, which is usually 30GB, and have no updates for 1 hour. Once the encoding is completed, EC shards are re-balanced; see [EC data balancing](#ec-data-balancing) below.
|
||||
|
||||
Note that if you have any collections, i.e. because you're using s3 where every bucket is a collection you explicitly need to specify the collection for the ec.encode command for them to be erasure coded `ec.encode -collection="collection" -fullPercent=95 -quietFor=1h`.
|
||||
|
||||
@@ -67,6 +67,14 @@ With servers added or removed, some data shards may not be laid out optimally. F
|
||||
|
||||
The default command is `ec.balance -force`. It will try to spread the data shards evenly to minimize the data shard loss risk.
|
||||
|
||||
EC shard re-balancing happens in three steps:
|
||||
|
||||
1. Duplicate EC shards for the same volume + server are deleted.
|
||||
2. EC shards are balanced across volumes for all racks.
|
||||
3. EC shards are balanced across volumes for individual racks.
|
||||
|
||||
Destination volume/racks for EC shards are selected based on capacity, favoring those with less preexisting shards in order to ensure an uniform distribution. Additionally, EC shards selection obey default [replica placement settings](Replication#the-meaning-of-replication-type) for the master server.
|
||||
|
||||
## How the read works?
|
||||
When all data shards are online, the read for one file key are assigned to one volume server (A) that has at least one data shard for the volume. Server A will read its copy of index file, and locate the volume server (B), and read from server B for the file key.
|
||||
|
||||
|
||||
Reference in new issue
Block a user