Files
seaweedfs/weed/storage/erasure_coding/ecbalancer/place.go
T
Chris Lu d4e39b499b EC placement: shared replica-placement resolver, snapshot + Place core, capacity fixes, tiering (#9621)
* Add shared super_block.ResolveReplicaPlacement; use it in ec_balance

* Add ecbalancer.FromActiveTopology snapshot constructor for EC encode/repair

* Add ecbalancer.Place greenfield/repair placement core (strict + durability-first)

* topology: add GetEffectiveAvailableEcShardSlots; FromActiveTopology uses shard-granular free slots

GetDisksWithEffectiveCapacity flattens reserved shard slots into volume slots via
integer truncation, so an in-flight EC task reserving a non-multiple-of-
DataShardsCount number of shards was lost from the snapshot and freeSlots was
over-reported. GetEffectiveAvailableEcShardSlots subtracts the full reservation
impact at shard granularity.

* ecbalancer.Place: reject nodes without a free disk of the requested type

FromActiveTopology keeps all disk types in the snapshot, so an SSD-only request
could be routed to a node with only HDD capacity (pickBestDiskOnNode then returns
disk 0 on the wrong tier). Filter rack/node selection to those with a free disk
of the requested type.

* ecbalancer.Place: enforce ReplicaPlacement DiffDataCenterCount (per-DC shard cap)

* ecbalancer: enforce DiffDataCenterCount in balance (cross-DC phase + cross-rack DC cap)

Adds a cross-DC corrective phase that drains data centers holding more than
DiffDataCenterCount shards of a volume, and a per-DC cap on cross-rack move
targets. Both are no-ops when DiffDataCenterCount is unset, so balance output is
unchanged for non-DC placements.

* topology: ratio-aware EC shard slots and provisional empty-disk slot

GetEffectiveAvailableEcShardSlots now takes the target collection's data-shard
count, so a 4+2 volume's larger shards are not over-counted at 10 per volume slot;
and it keeps the one provisional slot for freshly started empty servers that
report max=0, matching getEffectiveAvailableCapacityUnsafe. FromActiveTopology
threads the ratio through.

* ecbalancer.Place: explicit disk-type filter signal (fix HDD vs any ambiguity)

HardDriveType normalizes to "", which collided with "" meaning any disk. Add
Constraints.FilterDiskType and normalize both sides so a hdd request matches disks
reported as "" and never leaks to SSD, while filter=false still means any.

* ecbalancer: add clearShardAccounting for repair snapshot reconciliation

Clears one disk's copy of a shard from per-domain accounting and recomputes the
node-level union (preserving a kept copy on another disk of the same node), without
crediting capacity. Repair uses it to drop to-be-deleted copies before placing
missing shards.

* ecbalancer: don't cap cross-DC target racks when DiffRackCount is unset

len(racks)+1 wrongly limited each target rack (3 in a 2-rack cluster), so draining
a DC could stop short of the DiffDataCenterCount cap. Use MaxShardCount+1 as the
effectively-unlimited default.

* topology/ecbalancer: ratio-correct EC capacity accounting

Reservation shard slots (default ShardsPerVolumeSlot units) are now converted to
the target ratio before subtracting, and existing EC shards are charged by size
(targetDataShards/shardDataShards) so a 2+1 shard isn't counted as one 10+4 slot.
Per-shard ratio lookup is behind shardDataShards (OSS uses the standard ratio).

* ecbalancer.Place: candidate tiering and eligible-rack caps

Adds a per-disk eligibility/preference abstraction so Place supports:
- preferred-tag whole-plan retry (try disks carrying the earliest tags first,
  widen to all only if a tier cannot place every shard; reports
  SpilledOutsidePreferredTags),
- soft disk-type spill via DiskTypePolicy (Any/Prefer/Require): Prefer fills the
  preferred type then spills, reporting SpilledToOtherDiskType; Require filters,
- even per-rack caps that divide by racks holding an eligible disk, so a tiered
  cluster (e.g. SSDs in 2 of 4 racks) isn't capped impossibly low.
Disk tags carried via Node.AddDiskTags + FromActiveTopology.

* ecbalancer: export ClearShardAccounting for repair snapshot reconciliation

* ecbalancer: address review feedback (ratio rounding, bitmap walk, same-DC moves)

- topology/ecbalancer: round shard-reservation and existing-shard footprint up
  when converting to target-ratio shard slots, so a sub-slot reservation is not
  truncated to zero and free capacity is not overstated for low-data-shard
  layouts (targetDataShards < ds).
- erasure_coding: add ShardBits.All iterator and use it across the balancer,
  cross-DC phase, and placement scoring instead of scanning 0..MaxShardCount and
  probing Has on every id.
- ecbalancer: allow same-DC cross-rack moves when a DC already sits at its
  DiffDataCenterCount cap; a same-DC move leaves the DC total unchanged. Add a
  regression test that fails without the guard.
- ecbalancer cross-DC phase: pick targets via the eligible-aware
  pickNodeInRackEligible/pickBestDiskEligible helpers so the disk-type filter is
  honored and a 0 disk id is not mistaken for a valid selection.

* ecbalancer: test ecShardSlotsOnDisk fractional round-up

Cover the mixed-ratio path (targetDataShards < existing data shards) so a
shard's fractional footprint is never floored to zero and free capacity is not
overstated. Exercises the round-up via the targetDataShards parameter; OSS uses
the standard ratio at runtime while the enterprise build hits it with real
per-volume ratios.

* ecbalancer: assert node B rack in TestFromActiveTopology

* ecbalancer: split Destination into separate DataCenter and bare Rack

Replace the composite "dc:rack" Rack field on Destination with separate
DataCenter and bare Rack values, matching topology.DiskInfo and the worker-task
convention. Callers (and tests) read the data center directly instead of parsing
the composite with strings.SplitN.

* shell ec.balance: use utilization-based global balancing (parity with worker)

The shell's global rebalance phase balanced by raw shard count; switch it to
fractional fullness (shards/capacity), as the worker already does. On uniform
capacity the two agree; on heterogeneous capacity it fills nodes proportionally
instead of driving small-capacity nodes toward full.

Updates the heterogeneous-capacity regression test to assert even fullness
(~equal shards/capacity per node) rather than even shard count.

* ecbalancer: bounded-proportional per-DC shard spread

DiffDataCenterCount was enforced only as a ceiling (drain-to-cap), which could
leave a within-cap-but-lopsided DC distribution under a loose cap (e.g. 10/4 of 14
with cap=10). Now the cross-DC phase, the cross-rack DC guard, and Place all target
boundedMaxPerDC = min(DiffDataCenterCount, max(ceil(total/numDCs), parityShards)):
shards spread proportionally across DCs, but no tighter than the durability floor
(once each DC holds <= parityShards a DC loss is recoverable, so further spreading
only adds cross-DC/WAN traffic). No-op when DiffDataCenterCount is 0; identical to
before when the cap is the binding constraint.

* ecbalancer: drop DiffDataCenterCount enforcement for EC placement

The 1-byte volume ReplicaPlacement packs xyz into x*100+y*10+z<=255, so the DC
digit can only be 0-2 -- far too small to be a meaningful per-DC EC shard cap (a
cap of 1-2 would demand 7-14 DCs for a 10+4 volume). It's volume replica-placement,
not an EC spec. Removes the cross-DC balance phase, the DC guard in the cross-rack
phase, and the per-DC cap in Place (and the just-added bounded-proportional logic);
EC relies on the RP-independent rack/node even spread instead. Rack/node caps
(DiffRackCount/SameRackCount) are unchanged. Per-domain EC caps are left for a real
EC placement spec.

* ecbalancer: enforce per-disk durability cap; symmetric reserve/release

Place now refuses to put more than parityShards shards of a volume on a single
disk (pickBestDiskEligible skips a disk once it holds parityShards of the volume,
a hard cap not relaxed even in durability-first). Previously Place assigned by
free capacity, so a skewed near-full cluster could pile >parityShards onto one
disk -> losing it loses the volume; only distinct-disk count was checked. This
covers encode and repair (both route through Place); the caller skips/leaves the
volume rather than minting an unrecoverable layout.

Also makes reserveShard decrement freeSlots unconditionally, symmetric with
releaseShard's unconditional increment (the old guarded decrement could credit a
phantom slot on release if a shard were ever reserved onto a full disk).

* ecbalancer: add Topology.ReleaseVolumeShards (clear + credit) for greenfield encode

Releases all of a volume's shards from the snapshot and credits the freed disk
capacity, so a greenfield encode can plan as if stale EC shards from a prior failed
attempt are gone. Safe to credit because the encode task deletes stale shards
(cleanupStaleEcShards) before distributing the new ones. Distinct from
ClearShardAccounting (repair), which does not credit.

* ecbalancer: ReleaseVolumeShards credits node freeSlots, not just disks

releaseShard only increments per-disk freeSlots, but rack capacity is summed from
node freeSlots (buildRacks) and node freeSlots gates node eligibility. Crediting
only disks left a node/rack looking full after releasing stale shards, so a
greenfield encode still couldn't use the freed capacity. Now credits the node by
the total disk-slots freed.

* ecbalancer: correct PlacementMode docs (encode uses durability-first)

PlaceStrict was labeled '(encode)' but encode uses PlaceDurabilityFirst. Clarify
that durability-first is used by both encode and repair, reports relaxations in
PlaceResult.Relaxed, and never relaxes the per-disk durability cap.

* ecbalancer: treat SameRackCount as a direct per-node shard cap

The 3rd ReplicaPlacement digit now caps shards per node at exactly the digit
value, matching how DiffRackCount (2nd digit) caps per rack, instead of allowing
digit+1 per node. This makes the per-rack and per-node caps consistent and
matches the documented "digits cap EC shards per rack and per node" semantics;
e.g. 011 now means at most one shard per rack and one per node.
2026-05-22 20:22:09 -07:00

559 lines
19 KiB
Go

package ecbalancer
import (
"fmt"
"sort"
"strings"
"github.com/seaweedfs/seaweedfs/weed/storage/erasure_coding"
"github.com/seaweedfs/seaweedfs/weed/storage/super_block"
storagetypes "github.com/seaweedfs/seaweedfs/weed/storage/types"
)
// Constraints configures a Place call. Ratio resolves a collection's
// (dataShards, parityShards); nil uses the standard scheme. ReplicaPlacement,
// when non-nil, caps shards per rack (DiffRackCount = max shards/rack) and per
// node within a rack (SameRackCount = max shards/node); both digits are direct
// hard caps. The data-center digit (DiffDataCenterCount) is not honored:
// the 1-byte volume ReplicaPlacement can only encode 0-2 there, too small to be a
// meaningful per-DC EC shard cap, so EC relies on the rack/node even spread instead.
//
// DiskTypePolicy controls how DiskType constrains placement (Any / Prefer /
// Require). PreferredTags drives whole-plan tag tiering: Place tries disks
// carrying the earliest tags first and widens to all disks only if a tier cannot
// place every shard.
type Constraints struct {
DiskType string
DiskTypePolicy DiskTypePolicy
PreferredTags []string
ReplicaPlacement *super_block.ReplicaPlacement
Ratio func(collection string) (dataShards, parityShards int)
}
// DiskTypePolicy controls how Constraints.DiskType constrains placement.
type DiskTypePolicy int
const (
DiskTypeAny DiskTypePolicy = iota // any disk type
DiskTypePrefer // prefer DiskType, spill to other types if needed
DiskTypeRequire // only DiskType (HardDriveType when "")
)
// diskTypeEqual compares disk types after normalization, so "" and "hdd" (both
// HardDriveType) are equal.
func diskTypeEqual(a, b string) bool {
return storagetypes.ToDiskType(a).String() == storagetypes.ToDiskType(b).String()
}
// diskHasAnyTag reports whether the disk carries any of the given tags.
func diskHasAnyTag(d *disk, tags []string) bool {
for _, want := range tags {
for _, have := range d.tags {
if have == want {
return true
}
}
}
return false
}
// Destination is a chosen target for one shard. DataCenter and Rack are kept as
// separate values (matching topology.DiskInfo) rather than a "dc:rack" composite,
// so callers read them directly instead of parsing.
type Destination struct {
Node string
DiskID uint32
DataCenter string
Rack string // bare rack id within DataCenter
}
// PlaceResult holds the chosen destinations, which constraints had to be relaxed
// (durability-first only), and whether placement spilled outside the preferred
// disk type or tag tiers (for parity with today's logging).
type PlaceResult struct {
Destinations map[int]Destination
Relaxed []string
SpilledToOtherDiskType bool
SpilledOutsidePreferredTags bool
}
// PlacementMode selects the strictness/relaxation policy.
type PlacementMode int
const (
// PlaceStrict: caps and ReplicaPlacement are hard. Place fails rather than
// violate them, so the caller can defer (leave the volume as-is and retry).
PlaceStrict PlacementMode = iota
// PlaceDurabilityFirst (used by both encode and repair): relax per-type caps ->
// data/parity anti-affinity -> ReplicaPlacement, in that order, until each shard
// lands, reporting what was relaxed in PlaceResult.Relaxed. The per-disk
// durability cap (<= parityShards per disk) is never relaxed. Fails only if no
// disk has free capacity. Encode places best-effort this way and rebalancing
// tightens the spread afterward.
PlaceDurabilityFirst
)
// relaxation controls which placement-quality constraints are enforced on an
// attempt. preferring fresh nodes (repair's "avoid surviving-shard nodes") is not
// listed: pickNodeInRack already selects the node with the fewest shards of the
// volume, so survivors are deprioritized with built-in fallback.
type relaxation struct {
caps bool
antiAffinity bool
rp bool
}
func (r relaxation) relaxedNames() []string {
var n []string
if !r.caps {
n = append(n, "caps")
}
if !r.antiAffinity {
n = append(n, "anti-affinity")
}
if !r.rp {
n = append(n, "replica-placement")
}
return n
}
var strictAttempts = []relaxation{{caps: true, antiAffinity: true, rp: true}}
var durabilityAttempts = []relaxation{
{caps: true, antiAffinity: true, rp: true},
{caps: false, antiAffinity: true, rp: true},
{caps: false, antiAffinity: false, rp: true},
{caps: false, antiAffinity: false, rp: false},
}
type placedEntry struct {
node *Node
sid int
rackKey string
}
// Place assigns destinations for the `need` shard ids of volume (collection,vid),
// reading the volume's already-placed shards from the snapshot (so encode passes
// an empty-for-this-volume snapshot, repair passes one seeded with the surviving
// shards).
//
// Tag tiering (whole-plan retry): it tries the preferred-tag tiers in order, each
// a complete candidate set, and returns the first tier that places every shard;
// only when it falls through to the no-tag tier does it set
// SpilledOutsidePreferredTags. Within a tier, disk-type Prefer spills to other
// types per shard (SpilledToOtherDiskType); Require filters strictly.
func (t *Topology) Place(vid uint32, collection string, need []int, c Constraints, mode PlacementMode) (*PlaceResult, error) {
if len(need) == 0 {
return &PlaceResult{Destinations: map[int]Destination{}}, nil
}
vk := volKey{collection: collection, vid: vid}
dataShards, parityShards := erasure_coding.DataShardsCount, erasure_coding.ParityShardsCount
if c.Ratio != nil {
if d, p := c.Ratio(collection); d > 0 && p > 0 {
dataShards, parityShards = d, p
}
}
racks := buildRacks(t.nodes)
if len(racks) == 0 {
return nil, fmt.Errorf("no racks available for EC placement")
}
rackKeys := sortedKeys(racks)
// Disk-type eligibility (Require filters; Any/Prefer admit all) and the soft
// type preference applied in scoring under Prefer.
typeEligible := func(d *disk) bool {
if c.DiskTypePolicy == DiskTypeRequire {
return diskTypeEqual(d.diskType, c.DiskType)
}
return true
}
var prefer func(*disk) bool
if c.DiskTypePolicy == DiskTypePrefer {
prefer = func(d *disk) bool { return diskTypeEqual(d.diskType, c.DiskType) }
}
// Whole-plan retry over preferred-tag tiers; the first tier that places every
// shard wins. Reaching the no-tag tier means we spilled outside the tags.
tiers := tagTiers(c.PreferredTags)
var lastErr error
for _, tierTags := range tiers {
tt := tierTags
eligible := func(d *disk) bool {
return typeEligible(d) && (len(tt) == 0 || diskHasAnyTag(d, tt))
}
res, err := t.tryPlace(vk, need, dataShards, parityShards, racks, rackKeys, mode, c.ReplicaPlacement, eligible, prefer)
if err != nil {
lastErr = err
continue
}
if len(c.PreferredTags) > 0 && len(tierTags) == 0 {
res.SpilledOutsidePreferredTags = true
}
return res, nil
}
return nil, lastErr
}
// tagTiers returns the eligibility tag-sets in increasing breadth, ending with an
// empty set ("any disk"). Empty preferredTags yields a single any-disk tier.
func tagTiers(preferredTags []string) [][]string {
if len(preferredTags) == 0 {
return [][]string{nil}
}
tiers := make([][]string, 0, len(preferredTags)+1)
for k := range preferredTags {
tiers = append(tiers, append([]string(nil), preferredTags[:k+1]...))
}
return append(tiers, nil)
}
// tryPlace runs one whole-plan placement attempt restricted to disks satisfying
// `eligible`, with `prefer` (may be nil) ranking soft-preferred disks first. It
// journals reservations and rolls them all back if any shard cannot be placed, so
// a failed tier leaves the snapshot unchanged for the next attempt.
func (t *Topology) tryPlace(vk volKey, need []int, dataShards, parityShards int, racks map[string]*rack, rackKeys []string, mode PlacementMode, rp *super_block.ReplicaPlacement, eligible func(*disk) bool, prefer func(*disk) bool) (*PlaceResult, error) {
result := &PlaceResult{Destinations: make(map[int]Destination, len(need))}
// Per-type shard ids per rack (even caps), total shard count per rack
// (DiffRackCount), and the racks bearing each type (anti-affinity) — all seeded
// from the volume's existing shards.
shardsPerRack := map[bool]map[string][]int{true: {}, false: {}}
rackShardCount := map[string]int{}
bearing := map[bool]map[string]bool{true: {}, false: {}}
for _, n := range t.nodes {
info, ok := n.shards[vk]
if !ok {
continue
}
for sid := range info.shardBits.All() {
s := int(sid)
isData := s < dataShards
shardsPerRack[isData][n.rack] = append(shardsPerRack[isData][n.rack], s)
rackShardCount[n.rack]++
bearing[isData][n.rack] = true
}
}
// Even per-rack caps divide by racks that actually have an eligible free disk,
// not all racks (the snapshot keeps every disk type/tag), so a valid tiered
// cluster — e.g. SSDs in only 2 of 4 racks — is not capped impossibly low.
numEligibleRacks := 0
for _, rk := range rackKeys {
if rackHasFreeDisk(racks[rk], eligible) {
numEligibleRacks++
}
}
if numEligibleRacks < 1 {
numEligibleRacks = 1
}
attempts := strictAttempts
if mode == PlaceDurabilityFirst {
attempts = durabilityAttempts
}
var journal []placedEntry
relaxedSeen := map[string]bool{}
spilledType := false
placeShard := func(sid int, isData bool) bool {
typeTotal := dataShards
if !isData {
typeTotal = parityShards
}
for _, rl := range attempts {
node, diskID, spilled, ok := chooseShardDest(vk, sid, isData, dataShards, typeTotal, numEligibleRacks, parityShards, racks, rackKeys, rp, eligible, prefer, shardsPerRack[isData], rackShardCount, bearing, rl)
if !ok {
continue
}
reserveShard(node, vk, sid, diskID)
node.freeSlots--
racks[node.rack].freeSlots--
shardsPerRack[isData][node.rack] = append(shardsPerRack[isData][node.rack], sid)
rackShardCount[node.rack]++
bearing[isData][node.rack] = true
journal = append(journal, placedEntry{node: node, sid: sid, rackKey: node.rack})
result.Destinations[sid] = Destination{
Node: node.id,
DiskID: diskID,
DataCenter: node.dc,
Rack: strings.TrimPrefix(node.rack, node.dc+":"),
}
if spilled {
spilledType = true
}
for _, name := range rl.relaxedNames() {
relaxedSeen[name] = true
}
return true
}
return false
}
// Data shards first, then parity, so parity can avoid data-bearing racks.
for _, isData := range []bool{true, false} {
for _, sid := range shardsOfType(need, isData, dataShards) {
if placeShard(sid, isData) {
continue
}
for _, e := range journal {
releaseShard(e.node, vk, e.sid)
e.node.freeSlots++
racks[e.rackKey].freeSlots++
}
return nil, fmt.Errorf("cannot place EC shard %d of volume %d (collection %q)", sid, vk.vid, vk.collection)
}
}
result.SpilledToOtherDiskType = spilledType
for name := range relaxedSeen {
result.Relaxed = append(result.Relaxed, name)
}
sort.Strings(result.Relaxed)
return result, nil
}
// chooseShardDest selects a (node, disk) for one shard at the given relaxation
// level: pick a rack (even per-type cap + ReplicaPlacement caps + two-pass
// anti-affinity to the opposite type), then the least-loaded eligible node, then
// the best eligible disk. The third return reports whether the disk spilled off
// the soft-preferred type. ok=false when no rack/node/disk fits.
func chooseShardDest(vk volKey, sid int, isData bool, dataShards, typeTotal, numEligibleRacks, maxPerDisk int, racks map[string]*rack, rackKeys []string, rp *super_block.ReplicaPlacement, eligible func(*disk) bool, prefer func(*disk) bool, shardsPerRackType map[string][]int, rackShardCount map[string]int, bearing map[bool]map[string]bool, rl relaxation) (*Node, uint32, bool, bool) {
maxPerRack := numEligibleRacks*typeTotal + 1 // effectively unlimited when caps are relaxed
if rl.caps {
if maxPerRack = ceilDivide(typeTotal, numEligibleRacks); maxPerRack < 1 {
maxPerRack = 1
}
}
var anti map[string]bool
if rl.antiAffinity {
anti = bearing[!isData] // racks already holding the opposite shard type
}
if !rl.rp {
rp = nil
}
// A rack is eligible only if it is under the per-rack shard cap (DiffRackCount),
// enforced only when set (and relaxed with rp).
withinLimit := func(r string) bool {
if rp == nil {
return true
}
if rp.DiffRackCount > 0 && rackShardCount[r] >= rp.DiffRackCount {
return false
}
return true
}
destRack, ok := pickTarget(rackKeys, shardsPerRackType, maxPerRack, anti,
func(r string) bool { return racks[r].freeSlots > 0 && rackHasFreeDisk(racks[r], eligible) },
withinLimit)
if !ok {
return nil, 0, false, false
}
node := pickNodeInRackEligible(racks[destRack], vk, rp, eligible)
if node == nil {
return nil, 0, false, false
}
diskID, ok, spilled := pickBestDiskEligible(node, vk, eligible, prefer, sid, dataShards, maxPerDisk)
if !ok {
return nil, 0, false, false
}
return node, diskID, spilled, true
}
// nodeHasFreeDisk reports whether the node has a free disk satisfying eligible.
func nodeHasFreeDisk(n *Node, eligible func(*disk) bool) bool {
for _, d := range n.disks {
if d.freeSlots > 0 && eligible(d) {
return true
}
}
return false
}
// rackHasFreeDisk reports whether any node in the rack has a free eligible disk.
func rackHasFreeDisk(r *rack, eligible func(*disk) bool) bool {
for _, n := range r.nodes {
if n.freeSlots > 0 && nodeHasFreeDisk(n, eligible) {
return true
}
}
return false
}
// pickNodeInRackEligible is pickNodeInRack restricted to nodes that have a free
// eligible disk. FromActiveTopology keeps all disk types/tags in the snapshot, so
// without this a node with free volume slots but no eligible disk could be chosen.
func pickNodeInRackEligible(r *rack, vk volKey, rp *super_block.ReplicaPlacement, eligible func(*disk) bool) *Node {
var best *Node
bestCount := -1
for _, id := range sortedNodeKeys(r.nodes) {
node := r.nodes[id]
if node.freeSlots <= 0 {
continue
}
if !nodeHasFreeDisk(node, eligible) {
continue
}
count := volumeShardCount(node, vk)
if rp != nil && rp.SameRackCount > 0 && count >= rp.SameRackCount {
continue
}
if best == nil || count < bestCount {
best, bestCount = node, count
}
}
return best
}
// pickBestDiskEligible chooses the best eligible disk on a node, ranking
// soft-preferred disks (prefer != nil && prefer(d)) ahead of others so disk-type
// Prefer uses the preferred type when available but spills otherwise. Returns the
// disk id, whether one was found, and whether the chosen disk spilled off the
// preferred type.
func pickBestDiskEligible(node *Node, vk volKey, eligible func(*disk) bool, prefer func(*disk) bool, shardID, dataShardCount, maxPerDisk int) (uint32, bool, bool) {
isDataShard := dataShardCount > 0 && shardID < dataShardCount
info := node.shards[vk]
var bestDiskID uint32
bestScore := -1
bestPreferred := false
for _, diskID := range sortedDiskKeys(node.disks) {
d := node.disks[diskID]
if !eligible(d) || d.freeSlots <= 0 {
continue
}
existingShards := 0
hasData := false
hasParity := false
if info != nil {
bits := info.diskShardBits[diskID]
existingShards = bits.Count()
if dataShardCount > 0 {
for sid := range bits.All() {
if int(sid) < dataShardCount {
hasData = true
} else {
hasParity = true
}
}
}
}
// Durability: never put more than maxPerDisk (parityShards) shards of this
// volume on one disk, or losing that disk would lose more than EC can
// recover. Hard cap, enforced even under durability-first relaxation.
if maxPerDisk > 0 && existingShards >= maxPerDisk {
continue
}
score := d.shardCount*10 + existingShards*100
if dataShardCount > 0 {
if isDataShard && hasParity {
score += 1000
} else if !isDataShard && hasData {
score += 1000
}
}
preferred := prefer == nil || prefer(d)
if !preferred {
score += 100000 // strongly deprioritize spilling to a non-preferred type
}
if bestScore == -1 || score < bestScore {
bestScore = score
bestDiskID = diskID
bestPreferred = preferred
}
}
if bestScore == -1 {
return 0, false, false
}
return bestDiskID, true, prefer != nil && !bestPreferred
}
// clearShardAccounting removes one shard copy of a volume from the snapshot's
// per-domain accounting (the volume's shard bits) WITHOUT crediting disk capacity.
// It clears only the given physical disk's bit, then recomputes the node-level
// union from the remaining disk bits, so a kept copy of the same shard on another
// disk of the same node still counts toward caps / ReplicaPlacement / anti-affinity.
//
// Repair uses this to drop the duplicate/mismatched copies it plans to delete
// before placing missing shards, so those copies do not inflate placement
// accounting. Capacity is deliberately NOT credited: the deletes run only after
// the rebuilt shards are distributed, so the slots are not free at plan time. This
// is distinct from releaseShard, which credits freeSlots and clears the union.
func clearShardAccounting(node *Node, vk volKey, shardID int, diskID uint32) {
info, ok := node.shards[vk]
if !ok {
return
}
sid := erasure_coding.ShardId(shardID)
if bits, ok := info.diskShardBits[diskID]; ok {
info.diskShardBits[diskID] = bits.Clear(sid)
}
var union erasure_coding.ShardBits
for _, b := range info.diskShardBits {
union |= b
}
info.shardBits = union
}
// ClearShardAccounting drops one shard copy of a volume from placement accounting
// without crediting capacity (see clearShardAccounting). Repair calls it for each
// copy it plans to delete before placing missing shards, so those copies do not
// inflate caps/RP/anti-affinity. No-op for an unknown node.
func (t *Topology) ClearShardAccounting(nodeID, collection string, vid uint32, shardID int, diskID uint32) {
n, ok := t.nodes[nodeID]
if !ok {
return
}
clearShardAccounting(n, volKey{collection: collection, vid: vid}, shardID, diskID)
}
// ReleaseVolumeShards removes every shard of a volume from the snapshot and
// credits the freed disk capacity. A greenfield encode calls this so any stale
// EC shards left by a prior failed attempt (which the encode task deletes before
// distributing the new shards) neither occupy capacity nor skew anti-affinity /
// per-disk caps during planning. Unlike repair's ClearShardAccounting, it credits
// freeSlots because the deletes run before the new writes.
func (t *Topology) ReleaseVolumeShards(collection string, vid uint32) {
vk := volKey{collection: collection, vid: vid}
for _, n := range t.nodes {
info, ok := n.shards[vk]
if !ok {
continue
}
// freed is the total disk-slots the volume occupies on this node (a shard may
// sit on more than one disk). releaseShard credits each disk's freeSlots;
// credit the node's freeSlots by the same total, since rack capacity is summed
// from node freeSlots (buildRacks) and node freeSlots gates node eligibility.
freed := 0
for _, bits := range info.diskShardBits {
freed += bits.Count()
}
sids := make([]int, 0, info.shardBits.Count())
for sid := range info.shardBits.All() {
sids = append(sids, int(sid))
}
for _, sid := range sids {
releaseShard(n, vk, sid)
}
n.freeSlots += freed
delete(n.shards, vk)
}
}
// shardsOfType returns the sorted subset of need that are data shards (id <
// dataShards) when isData, else the parity subset.
func shardsOfType(need []int, isData bool, dataShards int) []int {
var out []int
for _, s := range need {
if (s < dataShards) == isData {
out = append(out, s)
}
}
sort.Ints(out)
return out
}