Files
seaweedfs/weed/storage/store_ec.go
T
Chris Lu 88c873ecd4 ec: uniform shard block layout (#10932)
* ec: uniform shard block layout

An EC volume is striped as 1GiB blocks until less than one row remains, then
1MiB blocks, and consecutive blocks land on different shards. With ec.encode's
-fullPercent 95 against the 30GiB default limit, ~30% of every volume sits in
that 1MiB tail, so a 4MB filer chunk there is five stripes on five servers.

New encodes now use one block per shard, sized ceil(datSize/dataShards) rounded
up to 1MiB and recorded in the .vif (EcShardConfig.block_size, also carried by
the .ecsum manifest). A needle now maps to one shard unless it is larger than
the block or straddles a boundary. The chosen size equals the legacy layout's
padded shard length for every input, so shard sizes, capacity math, and the
shard-size credibility checks are unchanged; only the byte placement moved.

Reads, decode, and scrub resolve the block sizes from the volume's .vif;
absence keeps the legacy interpretation, so existing EC volumes read exactly as
before. Rebuild is layout-agnostic. weed fix -ecx recovers the layout from the
.vif, else the .ecsum sidecar, and with neither de-stripes under both candidate
layouts and keeps the one that indexes more valid needles.

Same change in the Rust volume server, which now also streams the encode in
256KB sub-batches like Go instead of allocating whole blocks, and computes the
large-row count as shardSize/largeBlock to match Go on exact multiples. On a
26MB fixture both encoders produce byte-identical shards, and a Go-written .vif
parses in Rust with the block size intact.

* ec: resolve the rust ecx rebuild through the recorded layout

The Rust rebuild path regenerated a lost .ecx by scanning the logical .dat
through a hand-rolled pure-1MiB striping, which was already wrong for legacy
volumes with large-block rows and is wrong for any uniform volume with a block
past 1MiB. Route the scan through locate_data with the .vif-recorded block
size, the same mapping the read path uses. Also seed the new tests' random
data instead of the deprecated global math/rand.Read.

* ec: fail the Rust ecx rebuild on any shard read error

A read error mid-scan published the entries collected so far as a
successful .ecx, and read_at's byte count was ignored so a legal short
read passed as complete — a truncated or failing shard could produce a
silently incomplete recovery index. Exact-read semantics in
read_from_data_shards, error propagation in the needle walk, and a
truncated-shard regression test.

* ec: fail the mount on an unreadable or malformed vif

Both servers silently fell back to the legacy layout when an existing
.vif could not be read or parsed. Every new encode records a positive
uniform block size there, so the fallback mounted the same shards with
legacy offset math and could return wrong data. Absent stays legal
(legacy volumes predate the sidecar), and a zero-byte stub still reads
as absent (Go's MaybeLoadVolumeInfo convention, now mirrored in Rust);
a present-but-unreadable or malformed .vif fails the mount instead.

* ec: bound the reconstruct fan-out of one needle's intervals

A degraded interval fans out a read to every reachable shard location, each
with a buffer the size of the interval. Reading a needle's intervals in
parallel multiplied that by the interval concurrency: a needle spanning 8
blocks could hold 8 x MaxShardCount remote reads and buffers at once, where
the sequential version peaked at MaxShardCount. Give each needle a single
reconstruct budget its intervals share, held for the buffer's lifetime, so
separate reads stay independent but one read cannot multiply its own
fan-out.

* ec: drop the duplicated shard-size formula

calculateExpectedShardSize reimplemented the padding rule that
UniformBlockSize already owns — TestUniformBlockSizeMatchesLegacyShardSize
asserts the two agree for every input — so a change to the rule would have
had to be made in both. Defer to the helper, keeping the historic answer for
an empty .dat.

* ec: resolve the shard block layout from whatever records it

Four places still answered the layout question by inference when a record of
it was available, or accepted an answer that was not one:

- A mount with no .vif defaulted to the legacy layout; the bitrot sidecar
  records the same config at encode time, so take it when present, as
  weed fix -ecx already does. The vif itself is now parsed once per mount
  rather than twice.
- The Rust ecx rebuild derived its row count from the padded shard extent,
  which under the legacy layout reads a shard that is an exact large-block
  multiple as one row too many. Pass the encode-time .dat size from the .vif
  and keep the extent as the fallback.
- weed fix -ecx read the block size outside the EC-config guard (collapsing
  the unknown sentinel into a definitive legacy), only wrote the recovered
  layout back when the .vif was absent rather than unusable, and broke a
  scan tie by candidate order instead of the documented reach.
- The uniform layout tripped writeDatFile's large-block ambiguity guard,
  which cannot apply when the large and small blocks are the same size.

* ec: give the index-recovery tests a parseable vif

The fixtures wrote the literal bytes "volinfo" as the source .vif and the
recovery copies it verbatim, so the receiving server then mounted the volume
from a .vif it could not parse. That used to pass by silently defaulting to
the legacy layout; a mount now refuses a vif it cannot read, which is what
the tests were exercising all along without meaning to.

* ec: validate the layout a vif records, not just its syntax

Review follow-ups on the mount-strictness change:

- A .vif can parse and still record a block size no encoder could have
  produced (negative, or not a whole number of small blocks). Both servers
  took it and mapped every read through it. ValidateBlockSize / the Rust
  mirror now refuse the mount, the same way an unparseable vif does; 0 stays
  valid as the legacy two-tier layout.
- The bitrot-sidecar fallback accepted parity_shards == 0 and summed the
  counts in their own width, so values near the ceiling wrapped past the
  MaxShardCount bound. Require both counts and sum in a wider type.
- weed fix -ecx treated a config with only DataShards > 0 as usable, so a
  half-written .vif suppressed the recovery paths AND survived the rewrite.
  Require a complete, in-range config before trusting it.
- Returning the vif-load error left the .ecx and .ecj descriptors open;
  repeated mount attempts on malformed metadata could exhaust them.

* ec: refuse to act on a layout the metadata does not establish

- The worker encode only logged a failed .vif write and skipped it in the
  distribution set, and treated the .ecsum write as best-effort. A worker
  whose disk filled after the much larger shards landed could still
  distribute, mount, verify shard inventory, and delete the source replicas —
  leaving holders with shards whose geometry nothing records. Both writes and
  both inclusions are encode success conditions now.
- A generation-matching .ecsum that disagreed with the .vif geometry only
  disabled checksums in Go, and in Rust was not compared at all, so
  protection stayed On while reads used the other layout. Both files record
  the layout their generation was encoded with, so a disagreement now fails
  the mount.

* ec: reject an invalid recorded block size in weed fix -ecx

A .vif with valid shard counts but a negative or unaligned block size was
marked usable: a positive invalid value pinned the scan to a geometry that
de-stripes to garbage, and a negative one ran the dual scan but left the
invalid .vif in place afterwards. Validate it with the same rule the mount
applies, and when it fails leave the layout unknown so the scan recovers it
and the file is rewritten.

* ec: validate the sidecar layout weed fix -ecx recovers from

The .ecsum fallback was taken on DataShards > 0 alone, so a CRC-valid
sidecar carrying the wrong generation, an incomplete ratio, or an unaligned
block size would pin the reconstruction to one incorrect uniform-layout
candidate instead of letting the dual scan decide. Require generation 0, a
complete in-range ratio, and a valid block size; anything less leaves the
layout unknown, which is the answer that still recovers by scanning.

* ec: let only a genuinely absent sidecar choose the legacy layout

With no .vif the bitrot sidecar is the only record of a volume's layout, and
the mount fallback read a failed load, an unusable config, or a sidecar
stamped for another generation as "assume legacy". A uniform generation-0
volume could therefore mount with legacy or another generation's geometry and
answer reads with the wrong bytes. Present-but-unusable now fails the mount;
only actual absence keeps the legacy defaults. Shared as
EcShardConfigFromSidecar so every caller reads the sidecar the same way.

* ec: treat a recorded-but-impossible layout as corruption, not as legacy

- A .vif whose ecShardConfig is PRESENT but records an impossible ratio was
  answered with the default 10+4 and the legacy block layout, in both
  languages. That reads a uniform volume's shards at the wrong offsets and
  returns the wrong bytes. Only an entirely absent config still means "this
  predates the record"; a present one that cannot be true fails the mount.
- The shard-count bound summed two uint32 counts as int, which wraps on a
  32-bit build: 0x7fffffff + 0x7fffffff lands at -2 and slips under
  MaxShardCount. ValidEcShardCounts sums in uint64, and every EC call site
  that checked a recorded ratio now goes through it.

* ec: rebuild on the geometry the sidecar records, and flag it when it disagrees

The rebuild RPC passes BackgroundECContext, so RebuildEcFiles resolves the
layout itself — and it resolved a missing or invalid .vif to the default 10+4
with the legacy block size. Two consequences: a 12+4 volume was reconstructed
through a 10+4 matrix, which produces wrong bytes and never regenerates
shards 14-15; and the chosen geometry then contradicted a valid uniform
sidecar, which loadRebuildSidecar reported as BitrotOff — silently skipping
the input and regenerated-shard checksum checks precisely when the volume had
already lost its metadata.

The layout now resolves from the bitrot sidecar (found across the server's
disks, not just beside the base name) before falling back to the defaults,
and a present-but-impossible ratio fails instead of being replaced. A sidecar
that contradicts the chosen geometry is BitrotInvalid, which the existing
unsafeIgnoreSidecar override still lets an operator push past.

* ec: let the Rust rebuild read metadata off a sibling disk

read_ec_shard_config searches only the location the rebuild writes into, so a
volume whose .vif or generation-0 .ecsum sits on another of the server's
disks resolved to the default 10+4 with the legacy block layout — the Rust
half of the geometry-guessing the Go rebuild just stopped doing. It then
reconstructs a custom-ratio or uniform volume through the wrong
Reed-Solomon matrix and de-striping geometry.

The rebuild now looks for the .vif in its own location and then each sibling,
falls back to the generation-0 sidecar wherever that lives, and only defaults
when neither exists anywhere. The encode-time .dat size the ecx rebuild needs
is resolved the same way.

* ec: resolve a rebuild's vif from every directory that may hold it

RebuildEcFiles probed only <data-base>.vif. The caller knows the selected
location's index directory and the sibling locations, but passed neither for
metadata: additionalDirs carried shard directories only, and were searched
for shards and the checksum sidecar. A split -dir/-dir.idx layout, or a disk
holding only shards, therefore resolved a pre-sidecar custom-ratio volume to
10+4 and reconstructed through the wrong matrix — never regenerating shards
14-15.

The caller now hands over the index and sibling directories, and the resolver
probes the vif across all of them, matching what the Rust resolver already
does for both the vif and the sidecar.

* ec: make every rebuild consumer agree on the layout it resolved

- The post-rebuild bitrot backfill re-derived the geometry from this
  directory's .vif alone and dropped the block size entirely, so a rebuild
  that resolved its layout from a sibling, the sidecar, or a uniform vif wrote
  a manifest describing a DIFFERENT layout — one later mounts reject, or that
  covers only the default shard count. The layout is resolved once now,
  through an exported ResolveRebuildECContext, and the rebuild and the
  backfill share that answer.
- The Rust rebuild collected only each location's data directory, so a
  sibling's INDEX directory — where a split -dir/-dir.idx layout keeps
  .ecx/.ecj/.vif — was never probed, and a custom-ratio volume still resolved
  to 10+4 with the legacy layout. Both directories of every location are
  carried now, deduped against the rebuild's own.
- A shard delivery can bring the checksum manifest with it, but the receive
  path only writes the file: a server that already had the volume mounted kept
  its resolved protection state (off) until a remount. The mount RPC
  re-resolves it once the shards it describes have been added.

* ec: cover the rebuild's directory search with tests

Reviewers flagged the sibling index directory twice, and the fix that
closed it had no test of its own: the assembly sat inline in the rebuild
handler, reachable only through a gRPC call against a populated store.
Lifting it into rebuildSearchDirs / select_rebuild_location makes the
rule assertable — a sibling contributes BOTH its data and its index
directory, a shared index directory is listed once, and the rebuild's own
data directory never repeats.

Writing the Rust cases surfaced that the two implementations do not agree
on where the rebuild's own index directory belongs, and both are right:
Go's resolver takes a single directory list, so that directory has to be
inside it, while Rust's takes the rebuild's data and index directories as
their own arguments and would search them twice. The tests now state
which contract each side is holding to, so neither drifts into the
other's shape.

Pure refactor otherwise; no behaviour change.

* ec: search the index directory for the layout sidecar

The Rust resolver looked for the generation-0 .ecsum in the rebuild's
data directory and the sibling list, but not in the rebuild's own index
directory — while the .vif lookup directly above it did, and Go's
findBitrotSidecar has always checked both bases. On a split -dir/-dir.idx
location that directory is where the metadata lives, and callers leave it
out of the sibling list precisely because it is passed here separately,
so nothing searched it.

With no .vif anywhere the sidecar is the only surviving record of the
layout. Missing it resolved a 12+4 uniform volume to 10+4 with the legacy
striping — the test added here fails with (10, 4, 0) against the old
code — and the rebuild then reconstructs through the wrong matrix and
writes .ecx offsets that no reader can follow.

* ec: let the rebuild see its own index directory

The Rust rebuild takes a single flat directory list — the shape Go's
RebuildEcFiles uses — so it cannot be handed the rebuild location's index
directory separately the way the layout resolvers are, and the handler
was passing the sibling list, which deliberately omits exactly that
directory. On a split -dir/-dir.idx location that is where .ecx and .vif
live, so the shard and index lookups could not see them.

Go has always carried that directory in additionalDirs; this lines the
two call sites up.

* ec: let a config-free vif fall through to the layout sidecar

A .vif that carries no ecShardConfig answers nothing about the layout, so
it is no more informative than an absent one — but both trees treated its
mere existence as the end of the search. Go went straight to the 10+4
legacy defaults without consulting the sidecar at all; Rust returned
whatever ec_shard_config_from could make of a single directory. A 12+4
uniform volume with a legacy config-free vif therefore resolved as 10+4
legacy, and every read landed at the wrong shard offset.

The sidecar lookup was also single-directory on both sides, while a split
-dir/-dir.idx layout keeps .vif and .ecsum with the INDEX. Go's
findBitrotSidecar has always taken both bases; the callers here passed
only the data base, and the Rust bitrot resolver derived its path from
the data base alone. Rust's layout resolver now takes a candidate
directory list — data, index, then any siblings — and searches all of it,
which also removes the early return that made the vif's presence
decisive.

load_vif_info_across_dirs reported `dir` even when load_vif_info had
found the vif in `dir_idx`. Nothing reads that field today, so this
changes no behaviour; it stops the next caller that resolves the rest of
the volume's metadata against the answer from being sent to a disk
holding none of it.

Absence stays legal throughout: a volume with neither record is genuinely
legacy. Present-but-unusable still fails the mount, now in the
config-free-vif branch too.

* ec: activate a delivered sidecar on every per-disk runtime

A vid mounts as one EcVolume per disk, each with its own resolved
protection state, but the post-delivery reload used the first-match
lookup and so touched exactly one of them. The siblings kept reporting no
protection until a remount — and since shard distribution deduplicates
the metadata files onto the first target disk for a node, the runtime
that got the .ecsum is not necessarily the one the lookup returns.

Iterate every runtime instead, via a new FindAllEcVolumes and its Rust
mut equivalent. Combined with each runtime now resolving its sidecar
against its index directory as well as its data directory, a server
sharing one -dir.idx across its disks activates all of them from the
single delivered copy.

The Rust volume server had no post-mount reload at all; it gets one here,
matching Go.

* ec: resolve the delivered sidecar across every EC metadata directory

Reloading every per-disk runtime, added last round, did not by itself
make the delivered manifest reachable. Startup mirroring copies
.ecx/.ecj/.vif to every shard-bearing disk so each mounts
self-contained, but deliberately not .ecsum, and a repair delivers
exactly one copy. Each runtime was resolving against its own two
directories, so every sibling of the disk that received the file kept
reporting no protection however often it reloaded.

Resolve one authoritative copy across every EC metadata directory
instead of duplicating the file. Mirroring .ecsum would have to keep
pace with a file that is rewritten as shards are repaired, and would not
help the reported case at all: the delivery happens at runtime, and
mirroring only runs at startup.

The regression test pins both halves — a reload restricted to the
volume's own directories still finds nothing, and the same reload
given the server's metadata directories turns protection on.

* ec: ask every directory before writing a TOFU baseline

After a rebuild the opportunistic backfill asks whether this volume
already has a checksum manifest, and answered from the data base alone.
A split -dir/-dir.idx layout keeps the sidecar with the index, and a
multi-disk server may keep it on a sibling, so an existing manifest read
as absent.

The consequence is worse than a missed read. On a false "no" the backfill
writes a fresh sidecar at the data base from whatever the shards say right
now — and the data base is the first candidate every resolver checks, so
that TOFU baseline shadows the real manifest rather than sitting beside
it. A shard that was silently corrupt gets blessed, and the record that
would have caught it stops being consulted.

FindBitrotSidecar exports the search the package already used internally,
so the question is asked of the data base, the index base and the sibling
disks — the same candidates the rebuild resolves its layout from.

* ec: refuse a shard block size no encoder could have produced

weed fix -ecx derived one from the raw shard extent, so a truncated or
partially copied shard wrote a .vif that NewEcVolume then permanently
refuses — the volume the tool was run to rescue could never mount again.
An extent that is not a whole number of small blocks cannot have come
from a uniform encode, so it is no longer offered as a candidate, and
nothing unvalidated reaches the .vif.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: derive the .vif's dat size and block size from one measurement

VolumeEcShardsGenerate stat'ed the .dat before the encode while
WriteEcFiles stat'ed it again to size the blocks. A write landing
between the two produced a .vif whose own two fields describe different
files. WriteEcFiles now leaves both on the context, and fills a
placeholder context in place so the caller can read them back.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: keep the source volume until every holder serves its shard layout

The uniform layout rides in a .vif field older volume servers never
knew: they discard it, mount the shards as legacy and return wrong bytes
with nothing erroring, and the shard files are the same length either
way so no other check notices. The upgrade order lived only in the
release note. VolumeEcShardsInfo now reports the block size the holder
actually serves, in both the Go and Rust servers, and the pre-delete
verification refuses to drop the source unless every reachable holder
echoes the one the shards were encoded with — while a rollback still
exists. A server that predates the field answers 0, which is the
negative answer.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: drop the rebuild's dead block-size parameters

generateMissingEcFiles never reads largeBlockSize/smallBlockSize —
Reed-Solomon reconstruction is layout-agnostic — so passing the legacy
constants only advertised a layout the rebuild does not use. Also move
UniformBlockSize's doc off ValidateBlockSize.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: warn about EC defaults only when the mount used them

The "vif file not found, using defaults" warning fired even after the
bitrot sidecar supplied a non-default layout, sending anyone triaging
wrong bytes after the legacy layout the volume never mounted on.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: stat the distributed bitrot sidecar once

The strict check re-stat'ed the file immediately before the stat that
already gates inclusion, and a failed sidecar write now fails the encode
outright, so the first could only fire on a deletion between the two
lines.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: say what the reconstruct budget actually bounds

A shard's buffer stays in bufs until its interval reconstructs, which is
after the read that filled it released its permit, so the semaphore
bounds round trips in flight and not retained bytes. Peak memory is the
intervals reconstructing at once times the shards each reaches times the
interval size.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* test: let the fake volume server report its delivered EC layout

The pre-delete verification now asks each holder which shard block
layout it serves, and a fake that always answered "unset" looked exactly
like a volume server too old to know the field. Distribution ships the
.vif to every holder alongside its shards, so read the layout back out
of it as a real holder does.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7
2026-08-28 20:46:59 -07:00

993 lines
39 KiB
Go

package storage
import (
"context"
"errors"
"fmt"
"io"
"os"
"slices"
"sync"
"sync/atomic"
"time"
"github.com/klauspost/reedsolomon"
"golang.org/x/sync/semaphore"
"github.com/seaweedfs/seaweedfs/weed/glog"
"github.com/seaweedfs/seaweedfs/weed/operation"
"github.com/seaweedfs/seaweedfs/weed/pb"
"github.com/seaweedfs/seaweedfs/weed/pb/master_pb"
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
"github.com/seaweedfs/seaweedfs/weed/stats"
"github.com/seaweedfs/seaweedfs/weed/storage/erasure_coding"
"github.com/seaweedfs/seaweedfs/weed/storage/needle"
"github.com/seaweedfs/seaweedfs/weed/storage/types"
)
// errShardNotLocal indicates that the requested EC shard is simply not
// stored on this volume server. It is expected during normal reads when
// shards are spread across multiple servers, so callers should not log
// it as an error.
var errShardNotLocal = errors.New("ec shard not on this server")
// FindEcShardTargetLocation returns the disk that should receive a new
// shard / index file for (collection, vid). The selection order is:
//
// 1. a disk that already has the EC volume mounted (in-memory state),
// 2. a disk that owns the .ecx file on disk (volume not mounted yet),
// 3. any HDD with free space,
// 4. any disk with free space.
//
// Step 2 is the missing primitive that pinned subsequent shards to the
// first-shard disk during ec.rebuild. ec.rebuild only sets CopyEcxFile=true
// for the first shard, then relies on auto-select to land later shards on
// the same disk. Without an on-disk check, FindEcVolume returns nothing
// (no mount yet) and the fallback picks "any HDD with free space" — which
// can split shards from their index files across disks of the same node
// and lose them at startup. See issue #9212 and the orphan-shard
// reconciliation in #9244.
//
// dataShardCount is the data-shard count for this volume's EC layout (10
// for the OSS default, but custom ratios are supported via .vif). Callers
// pass it explicitly so this helper stays free of package-level constants
// — easier to mirror into builds that ship a different default ratio.
//
// Implementation walks s.Locations once and scores each disk by tier; the
// highest-tier disk wins, ties broken by free count. The earlier waterfall
// across four FindFreeLocation passes was equivalent but acquired
// volumesLock and ecVolumesLock RLocks (via VolumesLen / EcShardCount) up
// to four times per disk per call.
func (s *Store) FindEcShardTargetLocation(collection string, vid needle.VolumeId, dataShardCount int) *DiskLocation {
const (
tierAnyDisk = iota + 1
tierHDD
tierEcxOnDisk
tierMounted
)
var (
best *DiskLocation
bestTier int
bestFree int32
)
for _, loc := range s.Locations {
if loc.isDiskSpaceLow.Load() {
continue
}
freeCount := ecFreeShardCount(loc, dataShardCount)
if freeCount <= 0 {
continue
}
tier := tierAnyDisk
if loc.DiskType == types.HardDriveType {
tier = tierHDD
}
if loc.HasEcxFileOnDisk(collection, vid) {
tier = tierEcxOnDisk
}
if _, mounted := loc.FindEcVolume(vid); mounted {
tier = tierMounted
}
if best == nil || tier > bestTier || (tier == bestTier && freeCount > bestFree) {
best = loc
bestTier = tier
bestFree = freeCount
}
}
return best
}
// ecFreeShardCount returns the free EC shard capacity of loc, expressed
// in shard slots (not volume-equivalent slots). dataShardCount is the
// data-shard count of the EC layout being placed — see
// FindEcShardTargetLocation's docstring for why it's a parameter.
//
// FindFreeLocation in store.go does the same math but divides by
// DataShardsCount at the end. That truncation can exclude a disk that
// still has room for several individual shards (e.g. MaxVolumeCount=1,
// EcShardCount=1, dataShardCount=10 → reports 0 despite 9 free shard
// slots), which in this helper would re-route subsequent shards off the
// .ecx-owning disk and re-introduce the orphan-shard layout #9212 is
// trying to prevent. So we keep the result in shard slots throughout.
//
// MaxVolumeCount == 0 is the "unlimited" sentinel used elsewhere in the
// store (see hasFreeDiskLocation). Reporting a synthetic large free
// count keeps unlimited disks eligible while still letting tie-breaks
// prefer the less-loaded one.
func ecFreeShardCount(loc *DiskLocation, dataShardCount int) int32 {
if dataShardCount <= 0 {
return 0
}
if loc.MaxVolumeCount <= 0 {
const unlimitedFree = int32(1 << 30)
used := int32(loc.VolumesLen())*int32(dataShardCount) + int32(loc.EcShardCount())
if used >= unlimitedFree {
return 1
}
return unlimitedFree - used
}
free := (loc.MaxVolumeCount - int32(loc.VolumesLen())) * int32(dataShardCount)
free -= int32(loc.EcShardCount())
if free < 0 {
return 0
}
return free
}
func (s *Store) CollectErasureCodingHeartbeat() *master_pb.Heartbeat {
var ecShardMessages []*master_pb.VolumeEcShardInformationMessage
collectionEcShardSize := make(map[string]int64)
for diskId, location := range s.Locations {
location.ecVolumesLock.RLock()
for _, ecShards := range location.ecVolumes {
ecShardMessages = append(ecShardMessages, ecShards.ToVolumeEcShardInformationMessage(uint32(diskId))...)
for _, ecShard := range ecShards.Shards {
collectionEcShardSize[ecShards.Collection] += ecShard.Size()
}
}
location.ecVolumesLock.RUnlock()
}
for col, size := range collectionEcShardSize {
stats.VolumeServerDiskSizeGauge.WithLabelValues(col, "ec").Set(float64(size))
}
for col := range s.reportedEcCollections {
if _, stillHere := collectionEcShardSize[col]; !stillHere {
stats.VolumeServerDiskSizeGauge.DeleteLabelValues(col, "ec")
}
}
s.reportedEcCollections = make(map[string]struct{}, len(collectionEcShardSize))
for col := range collectionEcShardSize {
s.reportedEcCollections[col] = struct{}{}
}
return &master_pb.Heartbeat{
EcShards: ecShardMessages,
HasNoEcShards: len(ecShardMessages) == 0,
}
}
func (s *Store) MountEcShards(collection string, vid needle.VolumeId, shardId erasure_coding.ShardId, sourceDiskType string) error {
// The .ecx index file may live on a different disk than the one
// holding the .ec?? shard being mounted: ec.balance / ec.rebuild can
// place the .ecx on one local disk while later distributing shards
// across sibling disks of the same volume server. The per-disk
// IdxDirectory used by LoadEcShard would ENOENT the .ecx, so look up
// the .ecx owner across all DiskLocations once and route NewEcVolume
// at the directory that actually has the file. A 0-byte .ecx is
// treated as missing here (writeToFile can leave a stub on a failed
// EC distribute) so we still scan the rest of the disks.
ecxIdxDir, ecxFound := s.findEcxIdxDirForVolume(collection, vid)
// Collect failures so an all-disks-fail return reports every disk we
// tried rather than just the first one. Before this loop reordered
// itself to keep going after the first non-ENOENT error, a single
// shard-on-disk-without-.ecx situation would bail the loop and the
// operator saw "cannot open ec volume index" naming exactly one disk
// even when others held a valid index.
type diskError struct {
dir string
err error
}
var failures []diskError
for diskId, location := range s.Locations {
idxDir := location.IdxDirectory
if ecxFound {
// Fast path: if findEcxIdxDirForVolume already pointed at
// one of this disk's directories, the disk owns the .ecx
// and the local IdxDirectory is the right answer — skip
// the HasEcxFileOnDisk stat. Only fall back to the sibling
// disk's idxDir when this disk's directories are neither.
if location.IdxDirectory != ecxIdxDir && location.Directory != ecxIdxDir {
if !location.HasEcxFileOnDisk(collection, vid) {
idxDir = ecxIdxDir
}
}
}
ecVolume, err := location.loadEcShardWithIdxDir(collection, vid, shardId, idxDir)
if err == nil {
glog.V(0).Infof("MountEcShards %d.%d on disk ID %d", vid, shardId, diskId)
// Apply the orchestrator-supplied source disk type so the EC
// volume reports under it instead of the location's. Empty means
// "fall back to location's disk type" (#9423).
if sourceDiskType != "" {
ecVolume.SetDiskType(types.ToDiskType(sourceDiskType))
}
si := erasure_coding.NewShardsInfo()
si.Set(erasure_coding.NewShardInfo(shardId, erasure_coding.ShardSize(ecVolume.ShardSize())))
s.NewEcShardsChan <- &master_pb.VolumeEcShardInformationMessage{
Id: uint32(vid),
Collection: collection,
EcIndexBits: uint32(si.Bitmap()),
ShardSizes: si.SizesInt64(),
DiskType: string(ecVolume.DiskType()),
ExpireAtSec: ecVolume.ExpireAtSec,
DiskId: uint32(diskId),
EncodeTsNs: ecVolume.EncodeTsNs,
}
return nil
}
if errors.Is(err, os.ErrNotExist) {
// Shard or index not on this disk; another disk may own it.
continue
}
failures = append(failures, diskError{dir: location.Directory, err: err})
}
if len(failures) == 0 {
// No disk had the shard or the index; this volume server is not
// holding any artefacts for the requested shard. Name what we
// scanned for so the operator can tell "no .ecx anywhere" apart
// from "shard not on this server".
if !ecxFound {
return fmt.Errorf("MountEcShards %d.%d: no .ecx index found on any local disk", vid, shardId)
}
return fmt.Errorf("MountEcShards %d.%d not found on disk", vid, shardId)
}
// Some disks returned a real (non-ENOENT) error. Report them all so
// the caller can see whether the failures cluster around one disk
// (likely hardware) or are spread out (likely a config problem).
var b []byte
for i, f := range failures {
if i > 0 {
b = append(b, "; "...)
}
b = append(b, fmt.Sprintf("%s: %v", f.dir, f.err)...)
}
return fmt.Errorf("MountEcShards %d.%d load failures: %s", vid, shardId, string(b))
}
func (s *Store) UnmountEcShards(vid needle.VolumeId, shardId erasure_coding.ShardId, reqEncodeTsNs int64) error {
// Walk every disk: a split-disk reconciled volume can mount the same vid on
// more than one disk, so a first-match unmount would leave a sibling copy
// mounted and heartbeating. Emit one deletion delta per disk.
unmountedAny := false
var lastErr error
for diskId, location := range s.Locations {
ecShard, found := location.FindEcShard(vid, shardId)
if !found {
continue
}
// Capture the encode generation before unloading so the deletion delta
// carries it like the mount delta does.
var encodeTsNs int64
if ecVolume, ok := location.FindEcVolume(vid); ok {
encodeTsNs = ecVolume.EncodeTsNs
}
// Generation fence: when the caller carries a generation (stale-worker
// cleanup), only unmount a strictly-older generation; preserve a disk whose
// generation is same-or-newer, 0, or unknown, so a stale run cannot unmount
// a newer run's live shards. reqEncodeTsNs==0 (legacy/shell) unmounts all.
if reqEncodeTsNs > 0 && !(encodeTsNs > 0 && encodeTsNs < reqEncodeTsNs) {
glog.V(1).Infof("UnmountEcShards %d.%d disk_id:%d skipped: disk gen %d not older than request gen %d", vid, shardId, diskId, encodeTsNs, reqEncodeTsNs)
continue
}
if deleted := location.UnloadEcShard(vid, shardId); deleted {
si := erasure_coding.NewShardsInfo()
si.Set(erasure_coding.NewShardInfo(shardId, 0))
s.DeletedEcShardsChan <- &master_pb.VolumeEcShardInformationMessage{
Id: uint32(vid),
Collection: ecShard.Collection,
EcIndexBits: si.Bitmap(),
ShardSizes: si.SizesInt64(),
DiskType: string(ecShard.DiskType),
DiskId: uint32(diskId),
EncodeTsNs: encodeTsNs,
}
glog.V(0).Infof("UnmountEcShards %d.%d disk_id:%d", vid, shardId, diskId)
unmountedAny = true
} else {
lastErr = fmt.Errorf("UnmountEcShards %d.%d not found on disk %d", vid, shardId, diskId)
}
}
// nil when no disk held the shard (idempotent re-unmount).
if !unmountedAny {
return lastErr
}
return nil
}
func (s *Store) findEcShard(vid needle.VolumeId, shardId erasure_coding.ShardId) (diskId uint32, shard *erasure_coding.EcVolumeShard, found bool) {
for diskId, location := range s.Locations {
if v, found := location.FindEcShard(vid, shardId); found {
return uint32(diskId), v, found
}
}
return 0, nil, false
}
// FindEcShard returns the shard if any DiskLocation on this server holds it,
// along with that disk's id.
func (s *Store) FindEcShard(vid needle.VolumeId, shardId erasure_coding.ShardId) (diskId uint32, shard *erasure_coding.EcVolumeShard, found bool) {
return s.findEcShard(vid, shardId)
}
// FindEcVolumeWithShard returns the EcVolume on the disk that owns the given
// shard, plus the shard. The read guard must check the identity of the volume
// that owns the bytes served: on a multi-disk server one vid can hold shards
// from different encode runs across disks, so a first-match volume can differ.
func (s *Store) FindEcVolumeWithShard(vid needle.VolumeId, shardId erasure_coding.ShardId) (*erasure_coding.EcVolume, *erasure_coding.EcVolumeShard, bool) {
for _, location := range s.Locations {
if shard, found := location.FindEcShard(vid, shardId); found {
if ev, ok := location.FindEcVolume(vid); ok {
return ev, shard, true
}
}
}
return nil, nil, false
}
// EcMetadataDirs lists every directory on this server that could hold an EC
// volume's metadata — each disk's data and index directory. Startup mirroring
// gives each shard-bearing disk its own .ecx/.ecj/.vif, but the checksum
// sidecar is not mirrored and a repair delivers exactly one copy, so a runtime
// looking only at its own two directories cannot see it. Handing this list to
// the sidecar resolution keeps one authoritative copy reachable from every
// runtime rather than duplicating a file that is rewritten as shards are
// repaired and generations published.
func (s *Store) EcMetadataDirs() []string {
var dirs []string
appendDir := func(dir string) {
if dir == "" {
return
}
for _, existing := range dirs {
if existing == dir {
return
}
}
dirs = append(dirs, dir)
}
for _, location := range s.Locations {
appendDir(location.Directory)
appendDir(location.IdxDirectory)
}
return dirs
}
// FindAllEcVolumes returns every per-disk *EcVolume the store maps for vid. A vid
// can mount on N disks as N distinct runtimes, and the first-match FindEcVolume
// hides the siblings — so anything that has to reach the whole volume, rather
// than any one runtime of it, iterates this instead. Order mirrors the
// deterministic s.Locations order; nil when no disk holds the vid.
func (s *Store) FindAllEcVolumes(vid needle.VolumeId) []*erasure_coding.EcVolume {
var evs []*erasure_coding.EcVolume
for _, location := range s.Locations {
if ev, found := location.FindEcVolume(vid); found {
evs = append(evs, ev)
}
}
return evs
}
func (s *Store) FindEcVolume(vid needle.VolumeId) (*erasure_coding.EcVolume, bool) {
for _, location := range s.Locations {
if s, found := location.FindEcVolume(vid); found {
return s, true
}
}
return nil, false
}
// FindEcVolumeDiskIds returns every disk_id on this store that has an
// EcVolume entry for the given volume. Useful for diagnostic logging
// when a single FindEcVolume hit hides which disk is actually holding
// the mount (e.g., the ReceiveFile mounted-volume guard).
func (s *Store) FindEcVolumeDiskIds(vid needle.VolumeId) []uint32 {
var ids []uint32
for diskId, location := range s.Locations {
if _, found := location.FindEcVolume(vid); found {
ids = append(ids, uint32(diskId))
}
}
return ids
}
// shardFiles is a list of shard files, which is used to return the shard locations
func (s *Store) CollectEcShards(vid needle.VolumeId, shardFileNames []string) (ecVolume *erasure_coding.EcVolume, found bool) {
for _, location := range s.Locations {
if s, foundShards := location.CollectEcShards(vid, shardFileNames); foundShards {
ecVolume = s
found = true
}
}
return
}
func (s *Store) DestroyEcVolume(vid needle.VolumeId) {
for _, location := range s.Locations {
location.DestroyEcVolume(vid)
}
}
// UnloadEcVolume drops any in-memory EcVolume for vid from every disk and closes
// its fds without deleting files, so a following unlink frees the inodes.
func (s *Store) UnloadEcVolume(vid needle.VolumeId) {
for _, location := range s.Locations {
location.unloadEcVolume(vid)
}
}
func (s *Store) ReadEcShardNeedle(vid needle.VolumeId, n *needle.Needle, onReadSizeFn func(size types.Size)) (int, error) {
for _, location := range s.Locations {
if localEcVolume, found := location.FindEcVolume(vid); found {
offset, size, intervals, err := localEcVolume.LocateEcShardNeedle(n.Id, localEcVolume.Version)
if err != nil {
return 0, fmt.Errorf("locate in local ec volume: %w", err)
}
if size.IsDeleted() {
return 0, ErrorDeleted
}
if onReadSizeFn != nil {
onReadSizeFn(size)
}
glog.V(3).Infof("read ec volume %d offset %d size %d intervals:%+v", vid, offset.ToActualOffset(), size, intervals)
if len(intervals) > 1 {
glog.V(3).Infof("ReadEcShardNeedle needle id %s intervals:%+v", n.String(), intervals)
}
bytes, isDeleted, err := s.readEcShardIntervals(n.Id, localEcVolume, intervals)
// A holder reporting the needle deleted is authoritative -- deletes are
// never invented and never undone -- so answer that ahead of whatever
// error the shards it could not gather produced.
if isDeleted {
return 0, ErrorDeleted
}
if err != nil {
return 0, fmt.Errorf("ReadEcShardIntervals: %w", err)
}
err = n.ReadBytes(bytes, offset.ToActualOffset(), size, localEcVolume.Version)
if err != nil {
return 0, fmt.Errorf("ec volume %d needle %s offset %d size %d: %w", vid, n.String(), offset.ToActualOffset(), size, err)
}
return len(bytes), nil
}
}
return 0, fmt.Errorf("ec shard %d not found", vid)
}
var (
// ecRecoverBudget bounds the interval-sized buffers EC recovery holds across
// all concurrent reads: a peer that is slow to fail keeps a whole fan-out of
// them alive for the gRPC timeout, and a burst of those walked servers into an
// OOM. A var so a test can narrow it alongside ecRecoverSem.
ecRecoverBudget int64 = 256 << 20
ecRecoverSem = semaphore.NewWeighted(ecRecoverBudget)
)
// ecIntervalReadConcurrency bounds the fan-out of a single needle read. Blocks
// that follow each other in the .dat live on different shards, so a needle
// spanning several of them costs one round trip per block when read in sequence.
const ecIntervalReadConcurrency = 8
func (s *Store) readEcShardIntervals(needleId types.NeedleId, ecVolume *erasure_coding.EcVolume, intervals []erasure_coding.Interval) (data []byte, is_deleted bool, err error) {
if err = s.cachedLookupEcShardLocations(ecVolume); err != nil {
return nil, false, fmt.Errorf("failed to locate shard via master grpc %s: %v", s.MasterAddress, err)
}
var totalSize int
for _, interval := range intervals {
totalSize += int(interval.Size)
}
data = make([]byte, totalSize)
if len(intervals) <= 1 {
for _, interval := range intervals {
if is_deleted, err = s.readOneEcShardInterval(needleId, ecVolume, interval, data); err != nil {
return nil, is_deleted, err
}
}
return data, is_deleted, nil
}
errs := make([]error, len(intervals))
var deleted atomic.Bool
var wg sync.WaitGroup
sem := make(chan struct{}, ecIntervalReadConcurrency)
var pos int
for i, interval := range intervals {
buf := data[pos : pos+int(interval.Size)]
pos += int(interval.Size)
wg.Add(1)
sem <- struct{}{}
go func() {
defer wg.Done()
defer func() { <-sem }()
isDeleted, e := s.readOneEcShardInterval(needleId, ecVolume, interval, buf)
if isDeleted {
deleted.Store(true)
}
errs[i] = e
}()
}
wg.Wait()
is_deleted = deleted.Load()
for _, e := range errs {
if e != nil {
return nil, is_deleted, e
}
}
return data, is_deleted, nil
}
// readOneEcShardInterval fills data, which must be interval.Size long.
func (s *Store) readOneEcShardInterval(needleId types.NeedleId, ecVolume *erasure_coding.EcVolume, interval erasure_coding.Interval, data []byte) (is_deleted bool, err error) {
shardId, actualOffset := ecVolume.IntervalToShardIdAndOffset(interval)
// try local read
err = s.readLocalEcShardInterval(ecVolume, shardId, data, actualOffset)
if err == nil {
return
}
if errors.Is(err, errShardNotLocal) {
// expected when shards are spread across servers; fall through to remote read
glog.V(4).Infof("ec shard %d.%d not local, will try remote", ecVolume.VolumeId, shardId)
} else {
glog.V(0).Infof("read local ec shard %d.%d offset %d: %v", ecVolume.VolumeId, shardId, actualOffset, err)
}
ecVolume.ShardLocationsLock.RLock()
sourceDataNodes, hasShardIdLocation := ecVolume.ShardLocations[shardId]
ecVolume.ShardLocationsLock.RUnlock()
// try reading directly
if hasShardIdLocation {
_, is_deleted, err = s.readRemoteEcShardInterval(sourceDataNodes, needleId, ecVolume.VolumeId, shardId, data, actualOffset, ecVolume.EncodeTsNs)
if err == nil {
return
}
glog.V(0).Infof("read remote ec shard %d.%d locations: %v", ecVolume.VolumeId, shardId, err)
// Recovery below skips this very shard, so nothing else invalidates the
// location that just failed -- and a shard that has moved would otherwise
// be reconstructed on every read until the map's own window expires.
markShardLocationsStale(ecVolume)
}
// try reading by recovering from other shards
_, is_deleted, err = s.recoverOneRemoteEcShardInterval(needleId, ecVolume, shardId, data, actualOffset)
if err == nil {
return
}
glog.V(0).Infof("recover ec shard %d.%d : %v", ecVolume.VolumeId, shardId, err)
return
}
func forgetShardId(ecVolume *erasure_coding.EcVolume, shardId erasure_coding.ShardId) {
// failed to access the source data nodes, clear it up
ecVolume.ShardLocationsLock.Lock()
delete(ecVolume.ShardLocations, shardId)
ecVolume.ShardLocationsStale = true
ecVolume.ShardLocationsLock.Unlock()
}
// markShardLocationsStale flags the cached map for a prompt re-check after a
// read failed against one of its locations. Unlike forgetShardId it keeps the
// entry: a direct read is worth retrying, since a dead peer fails fast.
func markShardLocationsStale(ecVolume *erasure_coding.EcVolume) {
ecVolume.ShardLocationsLock.Lock()
ecVolume.ShardLocationsStale = true
ecVolume.ShardLocationsLock.Unlock()
}
// ecShardLocationsTTL is how long a cached shard map is trusted. A complete map
// is trusted longest. One short of DataShards, or one a failed read has just
// invalidated, is re-checked promptly: until it is, every read of the dropped
// shard skips the direct fetch and pays for a Reed-Solomon recovery instead.
func ecShardLocationsTTL(shardCount int, stale bool, ecCtx *erasure_coding.ECContext) time.Duration {
switch {
case stale || shardCount < ecCtx.DataShards:
return 11 * time.Second
case shardCount == ecCtx.Total():
return 37 * time.Minute
default:
return 7 * time.Minute
}
}
func (s *Store) cachedLookupEcShardLocations(ecVolume *erasure_coding.EcVolume) (err error) {
// Use the volume's own EC ratio so a custom-ratio volume (e.g. 9+3) is judged
// complete/recoverable against its real data-shard count, not the build default.
// In OSS the ratio is always 10+4, so this is a no-op.
ecCtx := ecVolume.ECContext
if ecCtx == nil {
ecCtx = erasure_coding.NewDefaultECContext(ecVolume.Collection, ecVolume.VolumeId)
}
// Judge the map and consume its mark in one critical section, so a mark raised
// from here on belongs to the next refresh rather than being cleared by this
// one -- the read that raised it has disproved the map this lookup is about to
// install. Recover goroutines mutate all three fields via forgetShardId, so an
// unguarded read would race a concurrent map write besides.
ecVolume.ShardLocationsLock.Lock()
stale := ecVolume.ShardLocationsStale
ttl := ecShardLocationsTTL(len(ecVolume.ShardLocations), stale, ecCtx)
fresh := ecVolume.ShardLocationsRefreshTime.Add(ttl).After(time.Now())
if !fresh {
ecVolume.ShardLocationsStale = false
}
ecVolume.ShardLocationsLock.Unlock()
if fresh {
return nil
}
glog.V(3).Infof("lookup and cache ec volume %d locations", ecVolume.VolumeId)
err = operation.WithMasterServerClient(context.Background(), false, s.MasterAddress, s.grpcDialOption, func(masterClient master_pb.SeaweedClient) error {
req := &master_pb.LookupEcVolumeRequest{
VolumeId: uint32(ecVolume.VolumeId),
}
resp, err := masterClient.LookupEcVolume(context.Background(), req)
if err != nil {
return fmt.Errorf("lookup ec volume %d: %v", ecVolume.VolumeId, err)
}
if len(resp.ShardIdLocations) < ecCtx.DataShards {
return fmt.Errorf("only %d shards found but %d required", len(resp.ShardIdLocations), ecCtx.DataShards)
}
ecVolume.ShardLocationsLock.Lock()
for _, shardIdLocations := range resp.ShardIdLocations {
shardId := erasure_coding.ShardId(shardIdLocations.ShardId)
delete(ecVolume.ShardLocations, shardId)
for _, loc := range shardIdLocations.Locations {
ecVolume.ShardLocations[shardId] = append(ecVolume.ShardLocations[shardId], pb.NewServerAddressFromLocation(loc))
}
}
ecVolume.ShardLocationsRefreshTime = time.Now()
ecVolume.ShardLocationsLock.Unlock()
return nil
})
if err != nil {
// The lookup this mark was consumed for did not answer, and the refresh time
// is only advanced on success -- so without putting the mark back, a map a
// read had already disproved would be trusted for its full window again.
// Marking unconditionally is safe: where no mark was consumed the refresh
// time is stale anyway, and the next call looks up whatever the mark says.
markShardLocationsStale(ecVolume)
}
return
}
func (s *Store) readLocalEcShardInterval(ecVolume *erasure_coding.EcVolume, shardId erasure_coding.ShardId, buf []byte, offset int64) error {
// Resolve the shard together with the EcVolume on the disk that owns it; the
// shard may live on a sibling disk of this server.
ownerVolume, shard, found := s.FindEcVolumeWithShard(ecVolume.VolumeId, shardId)
if !found {
return fmt.Errorf("shard %d for volume %d: %w", shardId, ecVolume.VolumeId, errShardNotLocal)
}
// Skip a local shard whose identity doesn't match the caller's index, so the
// read recovers from the correct generation. Lenient only when the caller has
// no identity (pre-upgrade): a known caller must not accept an unstamped local
// shard, which would serve a stale pre-upgrade generation.
if ecVolume.EncodeTsNs != 0 && ecVolume.EncodeTsNs != ownerVolume.EncodeTsNs {
glog.V(1).Infof("skip local ec shard %d.%d from a different encode run: caller EncodeTsNs %d, local %d", ecVolume.VolumeId, shardId, ecVolume.EncodeTsNs, ownerVolume.EncodeTsNs)
return fmt.Errorf("shard %d for volume %d: %w", shardId, ecVolume.VolumeId, errShardNotLocal)
}
readBytes, err := shard.ReadAt(buf, offset)
if err != nil {
return fmt.Errorf("failed to read local EC shard %d for volume %d: %v", shardId, ecVolume.VolumeId, err)
}
if got, want := readBytes, len(buf); got != want {
return fmt.Errorf("expected %d bytes for local EC shard %d on volume %d, got %d", want, shardId, ecVolume.VolumeId, got)
}
return nil
}
func (s *Store) readRemoteEcShardInterval(sourceDataNodes []pb.ServerAddress, needleId types.NeedleId, vid needle.VolumeId, shardId erasure_coding.ShardId, buf []byte, offset int64, expectedEncodeTsNs int64) (n int, is_deleted bool, err error) {
if len(sourceDataNodes) == 0 {
return 0, false, fmt.Errorf("failed to find ec shard %d.%d", vid, shardId)
}
for _, sourceDataNode := range sourceDataNodes {
glog.V(3).Infof("read remote ec shard %d.%d from %s", vid, shardId, sourceDataNode)
n, is_deleted, err = s.doReadRemoteEcShardInterval(sourceDataNode, needleId, vid, shardId, buf, offset, expectedEncodeTsNs)
if err == nil {
return
}
glog.V(1).Infof("read remote ec shard %d.%d from %s: %v", vid, shardId, sourceDataNode, err)
}
return
}
func (s *Store) doReadRemoteEcShardInterval(sourceDataNode pb.ServerAddress, needleId types.NeedleId, vid needle.VolumeId, shardId erasure_coding.ShardId, buf []byte, offset int64, expectedEncodeTsNs int64) (n int, is_deleted bool, err error) {
err = operation.WithVolumeServerClient(false, sourceDataNode, s.grpcDialOption, func(client volume_server_pb.VolumeServerClient) error {
// copy data slice
shardReadClient, err := client.VolumeEcShardRead(context.Background(), &volume_server_pb.VolumeEcShardReadRequest{
VolumeId: uint32(vid),
ShardId: uint32(shardId),
Offset: offset,
Size: int64(len(buf)),
FileKey: uint64(needleId),
EncodeTsNs: expectedEncodeTsNs,
})
if err != nil {
return fmt.Errorf("failed to start reading ec shard %d.%d from %s: %v", vid, shardId, sourceDataNode, err)
}
for {
resp, receiveErr := shardReadClient.Recv()
if receiveErr == io.EOF {
break
}
if receiveErr != nil {
return fmt.Errorf("receiving ec shard %d.%d from %s: %v", vid, shardId, sourceDataNode, receiveErr)
}
// Validate the served shard's identity client-side, so the guard holds
// even against a pre-upgrade server that ignored the request field (it
// returns 0). A mismatch fails the read; the caller recovers from parity.
if expectedEncodeTsNs != 0 && resp.EncodeTsNs != expectedEncodeTsNs {
return fmt.Errorf("ec shard %d.%d from %s belongs to a different encode run (want %d, got %d)", vid, shardId, sourceDataNode, expectedEncodeTsNs, resp.EncodeTsNs)
}
if resp.IsDeleted {
is_deleted = true
}
copy(buf[n:n+len(resp.Data)], resp.Data)
n += len(resp.Data)
}
return nil
})
if err != nil {
return 0, is_deleted, fmt.Errorf("read ec shard %d.%d from %s: %v", vid, shardId, sourceDataNode, err)
}
// A non-deleted interval must arrive whole: the server stamps EncodeTsNs only
// on chunks that carry bytes, so a short or empty stream (e.g. immediate EOF
// from a pre-upgrade or stale server) leaves the buffer partly zero-filled and
// unvalidated. Reject it so the caller recovers from parity. The is_deleted
// short-circuit legitimately returns n=0 with no data and is exempt, matching
// readLocalEcShardInterval's got==len(buf) rule for the local path.
if !is_deleted && n != len(buf) {
return n, is_deleted, fmt.Errorf("short read ec shard %d.%d from %s: got %d want %d", vid, shardId, sourceDataNode, n, len(buf))
}
return
}
// gatherEcShardIntervals collects the same interval from every shard but shardIdToRecover,
// returning one buffer per shard and nil where the shard could not be gathered. It stops
// once DataShards of them are in hand: reconstruction consumes no more than that.
func (s *Store) gatherEcShardIntervals(needleId types.NeedleId, ecVolume *erasure_coding.EcVolume, ecCtx *erasure_coding.ECContext, shardIdToRecover erasure_coding.ShardId, size int, offset int64) (shardIntervals [][]byte, isDeleted bool) {
shardIntervals = make([][]byte, ecCtx.Total())
// A shard this server already holds costs no round trip and no peer buffer,
// so seed those before asking peers for the rest.
available := 0
for shardId := erasure_coding.ShardId(0); int(shardId) < ecCtx.Total() && available < ecCtx.DataShards; shardId++ {
if shardId == shardIdToRecover {
continue
}
if _, _, found := s.FindEcVolumeWithShard(ecVolume.VolumeId, shardId); !found {
continue
}
data := make([]byte, size)
if localErr := s.readLocalEcShardInterval(ecVolume, shardId, data, offset); localErr != nil {
glog.V(3).Infof("recover: read local ec shard %d.%d: %v", ecVolume.VolumeId, shardId, localErr)
continue
}
shardIntervals[shardId] = data
available++
}
var candidates []erasure_coding.ShardId
candidateLocations := make(map[erasure_coding.ShardId][]pb.ServerAddress)
ecVolume.ShardLocationsLock.RLock()
for shardId, locations := range ecVolume.ShardLocations {
// skip the shard being recovered, one already seeded locally, or an empty shard
if shardId == shardIdToRecover || int(shardId) >= ecCtx.Total() || shardIntervals[shardId] != nil {
continue
}
if len(locations) == 0 {
glog.V(3).Infof("readRemoteEcShardInterval missing %d.%d from %+v", ecVolume.VolumeId, shardId, locations)
continue
}
candidates = append(candidates, shardId)
candidateLocations[shardId] = locations
}
ecVolume.ShardLocationsLock.RUnlock()
// The recover goroutines run concurrently, so the deleted flag is collected
// atomically and folded into the named return after they join, rather than each
// goroutine writing the shared bool directly.
var isDeletedFlag atomic.Bool
// Reconstruction consumes DataShards buffers, so reading every remaining shard
// holds a third more than that and asks a third more of peers that are, by
// then, already struggling. Widen only when a wave falls short.
for len(candidates) > 0 && available < ecCtx.DataShards {
wave := min(ecCtx.DataShards-available, len(candidates))
var wg sync.WaitGroup
var fetched atomic.Int64
for _, shardId := range candidates[:wave] {
locations := candidateLocations[shardId]
wg.Add(1)
go func() {
defer wg.Done()
data := make([]byte, size)
nRead, isDeleted, readErr := s.readRemoteEcShardInterval(locations, needleId, ecVolume.VolumeId, shardId, data, offset, ecVolume.EncodeTsNs)
if readErr != nil {
glog.V(3).Infof("recover: readRemoteEcShardInterval %d.%d %d bytes from %+v: %v", ecVolume.VolumeId, shardId, nRead, locations, readErr)
forgetShardId(ecVolume, shardId)
}
if isDeleted {
isDeletedFlag.Store(true)
}
if nRead == size {
shardIntervals[shardId] = data
fetched.Add(1)
}
}()
}
wg.Wait()
candidates = candidates[wave:]
available += int(fetched.Load())
if isDeletedFlag.Load() {
// every shard of a deleted needle answers deleted, so another wave cannot help
break
}
}
return shardIntervals, isDeletedFlag.Load()
}
// checkEcShardRebuildable rejects a parity target. ReconstructData rebuilds data shards
// only, so it leaves a parity slot nil and reports no error, and the caller would copy
// that out as a zero-filled buffer and report a successful read.
func checkEcShardRebuildable(ecVolume *erasure_coding.EcVolume, ecCtx *erasure_coding.ECContext, shardIdToRecover erasure_coding.ShardId) error {
if int(shardIdToRecover) >= ecCtx.DataShards {
return fmt.Errorf("cannot reconstruct shard %d.%d: only data shards can be rebuilt, %d of %d are parity",
ecVolume.VolumeId, shardIdToRecover, ecCtx.ParityShards, ecCtx.Total())
}
return nil
}
// reconstructEcShardInterval rebuilds one shard's interval in place from the others.
func reconstructEcShardInterval(ecVolume *erasure_coding.EcVolume, ecCtx *erasure_coding.ECContext, shardIntervals [][]byte, shardIdToRecover erasure_coding.ShardId) error {
if err := checkEcShardRebuildable(ecVolume, ecCtx, shardIdToRecover); err != nil {
return err
}
enc, err := reedsolomon.New(ecCtx.DataShards, ecCtx.ParityShards)
if err != nil {
return fmt.Errorf("failed to create encoder: %w", err)
}
// Count and log available shards for diagnostics
availableShards := make([]erasure_coding.ShardId, 0, ecCtx.Total())
missingShards := make([]erasure_coding.ShardId, 0, ecCtx.ParityShards+1)
for shardId := 0; shardId < ecCtx.Total(); shardId++ {
if shardIntervals[shardId] != nil {
availableShards = append(availableShards, erasure_coding.ShardId(shardId))
} else {
missingShards = append(missingShards, erasure_coding.ShardId(shardId))
}
}
glog.V(3).Infof("recover ec shard %d.%d: %d shards available %v, %d missing %v",
ecVolume.VolumeId, shardIdToRecover,
len(availableShards), availableShards,
len(missingShards), missingShards)
if len(availableShards) < ecCtx.DataShards {
return fmt.Errorf("cannot recover shard %d.%d: only %d shards available %v, need at least %d (missing: %v)",
ecVolume.VolumeId, shardIdToRecover,
len(availableShards), availableShards,
ecCtx.DataShards, missingShards)
}
// Rebuild only what was asked for. ReconstructData rebuilds every missing data
// shard, and a gather that stopped at DataShards can leave up to ParityShards of
// them missing -- an interval-sized buffer and a decode each, discarded unread.
// The mask is Total() long, not DataShards: reedsolomon documents both lengths
// but indexes the short one past its end when a parity shard is absent, which
// here it usually is.
required := make([]bool, ecCtx.Total())
required[shardIdToRecover] = true
if err := enc.ReconstructSome(shardIntervals, required); err != nil {
return fmt.Errorf("failed to reconstruct data for shard %d.%d with %d available shards %v: %w",
ecVolume.VolumeId, shardIdToRecover, len(availableShards), availableShards, err)
}
return nil
}
func (s *Store) recoverOneRemoteEcShardInterval(needleId types.NeedleId, ecVolume *erasure_coding.EcVolume, shardIdToRecover erasure_coding.ShardId, buf []byte, offset int64) (n int, is_deleted bool, err error) {
glog.V(3).Infof("recover ec shard %d.%d from other locations", ecVolume.VolumeId, shardIdToRecover)
// Reconstruct with the volume's OWN EC ratio (loaded from its .vif), not the
// build default, so a custom-ratio volume (e.g. 9+3) is decoded with the matrix
// that actually produced its shards -- decoding it as 10+4 would corrupt the
// recovered bytes. In OSS the ratio is always 10+4, so this is a no-op.
ecCtx := ecVolume.ECContext
if ecCtx == nil {
ecCtx = erasure_coding.NewDefaultECContext(ecVolume.Collection, ecVolume.VolumeId)
}
// checked before the gather: a doomed target should not cost a fan-out, nor drop a
// sibling's location through forgetShardId on the way to failing
if err := checkEcShardRebuildable(ecVolume, ecCtx, shardIdToRecover); err != nil {
return 0, false, err
}
// Charge the buffers this recovery is about to hold against the budget, so a
// burst of them queues here rather than on the heap: DataShards gathered, plus
// the one the rebuild allocates for the shard it recreates.
weight := int64(len(buf)) * int64(ecCtx.DataShards+1)
if weight > ecRecoverBudget {
// An interval whose fan-out outgrows the whole budget takes all of it and
// so runs alone, rather than blocking forever on an acquire that can never
// succeed. The cap is then one such recovery, not a burst of them.
weight = ecRecoverBudget
}
if err = ecRecoverSem.Acquire(context.Background(), weight); err != nil {
return 0, false, err
}
defer ecRecoverSem.Release(weight)
shardIntervals, is_deleted := s.gatherEcShardIntervals(needleId, ecVolume, ecCtx, shardIdToRecover, len(buf), offset)
if err := reconstructEcShardInterval(ecVolume, ecCtx, shardIntervals, shardIdToRecover); err != nil {
return 0, is_deleted, err
}
glog.V(4).Infof("recovered ec shard %d.%d from other locations", ecVolume.VolumeId, shardIdToRecover)
copy(buf, shardIntervals[shardIdToRecover])
return len(buf), is_deleted, nil
}
func (s *Store) EcVolumes() (ecVolumes []*erasure_coding.EcVolume) {
for _, location := range s.Locations {
location.ecVolumesLock.RLock()
for _, v := range location.ecVolumes {
ecVolumes = append(ecVolumes, v)
}
location.ecVolumesLock.RUnlock()
}
slices.SortFunc(ecVolumes, func(a, b *erasure_coding.EcVolume) int {
return int(a.VolumeId) - int(b.VolumeId)
})
return ecVolumes
}