mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-09-08 15:41:15 +02:00
* ec: uniform shard block layout An EC volume is striped as 1GiB blocks until less than one row remains, then 1MiB blocks, and consecutive blocks land on different shards. With ec.encode's -fullPercent 95 against the 30GiB default limit, ~30% of every volume sits in that 1MiB tail, so a 4MB filer chunk there is five stripes on five servers. New encodes now use one block per shard, sized ceil(datSize/dataShards) rounded up to 1MiB and recorded in the .vif (EcShardConfig.block_size, also carried by the .ecsum manifest). A needle now maps to one shard unless it is larger than the block or straddles a boundary. The chosen size equals the legacy layout's padded shard length for every input, so shard sizes, capacity math, and the shard-size credibility checks are unchanged; only the byte placement moved. Reads, decode, and scrub resolve the block sizes from the volume's .vif; absence keeps the legacy interpretation, so existing EC volumes read exactly as before. Rebuild is layout-agnostic. weed fix -ecx recovers the layout from the .vif, else the .ecsum sidecar, and with neither de-stripes under both candidate layouts and keeps the one that indexes more valid needles. Same change in the Rust volume server, which now also streams the encode in 256KB sub-batches like Go instead of allocating whole blocks, and computes the large-row count as shardSize/largeBlock to match Go on exact multiples. On a 26MB fixture both encoders produce byte-identical shards, and a Go-written .vif parses in Rust with the block size intact. * ec: resolve the rust ecx rebuild through the recorded layout The Rust rebuild path regenerated a lost .ecx by scanning the logical .dat through a hand-rolled pure-1MiB striping, which was already wrong for legacy volumes with large-block rows and is wrong for any uniform volume with a block past 1MiB. Route the scan through locate_data with the .vif-recorded block size, the same mapping the read path uses. Also seed the new tests' random data instead of the deprecated global math/rand.Read. * ec: fail the Rust ecx rebuild on any shard read error A read error mid-scan published the entries collected so far as a successful .ecx, and read_at's byte count was ignored so a legal short read passed as complete — a truncated or failing shard could produce a silently incomplete recovery index. Exact-read semantics in read_from_data_shards, error propagation in the needle walk, and a truncated-shard regression test. * ec: fail the mount on an unreadable or malformed vif Both servers silently fell back to the legacy layout when an existing .vif could not be read or parsed. Every new encode records a positive uniform block size there, so the fallback mounted the same shards with legacy offset math and could return wrong data. Absent stays legal (legacy volumes predate the sidecar), and a zero-byte stub still reads as absent (Go's MaybeLoadVolumeInfo convention, now mirrored in Rust); a present-but-unreadable or malformed .vif fails the mount instead. * ec: bound the reconstruct fan-out of one needle's intervals A degraded interval fans out a read to every reachable shard location, each with a buffer the size of the interval. Reading a needle's intervals in parallel multiplied that by the interval concurrency: a needle spanning 8 blocks could hold 8 x MaxShardCount remote reads and buffers at once, where the sequential version peaked at MaxShardCount. Give each needle a single reconstruct budget its intervals share, held for the buffer's lifetime, so separate reads stay independent but one read cannot multiply its own fan-out. * ec: drop the duplicated shard-size formula calculateExpectedShardSize reimplemented the padding rule that UniformBlockSize already owns — TestUniformBlockSizeMatchesLegacyShardSize asserts the two agree for every input — so a change to the rule would have had to be made in both. Defer to the helper, keeping the historic answer for an empty .dat. * ec: resolve the shard block layout from whatever records it Four places still answered the layout question by inference when a record of it was available, or accepted an answer that was not one: - A mount with no .vif defaulted to the legacy layout; the bitrot sidecar records the same config at encode time, so take it when present, as weed fix -ecx already does. The vif itself is now parsed once per mount rather than twice. - The Rust ecx rebuild derived its row count from the padded shard extent, which under the legacy layout reads a shard that is an exact large-block multiple as one row too many. Pass the encode-time .dat size from the .vif and keep the extent as the fallback. - weed fix -ecx read the block size outside the EC-config guard (collapsing the unknown sentinel into a definitive legacy), only wrote the recovered layout back when the .vif was absent rather than unusable, and broke a scan tie by candidate order instead of the documented reach. - The uniform layout tripped writeDatFile's large-block ambiguity guard, which cannot apply when the large and small blocks are the same size. * ec: give the index-recovery tests a parseable vif The fixtures wrote the literal bytes "volinfo" as the source .vif and the recovery copies it verbatim, so the receiving server then mounted the volume from a .vif it could not parse. That used to pass by silently defaulting to the legacy layout; a mount now refuses a vif it cannot read, which is what the tests were exercising all along without meaning to. * ec: validate the layout a vif records, not just its syntax Review follow-ups on the mount-strictness change: - A .vif can parse and still record a block size no encoder could have produced (negative, or not a whole number of small blocks). Both servers took it and mapped every read through it. ValidateBlockSize / the Rust mirror now refuse the mount, the same way an unparseable vif does; 0 stays valid as the legacy two-tier layout. - The bitrot-sidecar fallback accepted parity_shards == 0 and summed the counts in their own width, so values near the ceiling wrapped past the MaxShardCount bound. Require both counts and sum in a wider type. - weed fix -ecx treated a config with only DataShards > 0 as usable, so a half-written .vif suppressed the recovery paths AND survived the rewrite. Require a complete, in-range config before trusting it. - Returning the vif-load error left the .ecx and .ecj descriptors open; repeated mount attempts on malformed metadata could exhaust them. * ec: refuse to act on a layout the metadata does not establish - The worker encode only logged a failed .vif write and skipped it in the distribution set, and treated the .ecsum write as best-effort. A worker whose disk filled after the much larger shards landed could still distribute, mount, verify shard inventory, and delete the source replicas — leaving holders with shards whose geometry nothing records. Both writes and both inclusions are encode success conditions now. - A generation-matching .ecsum that disagreed with the .vif geometry only disabled checksums in Go, and in Rust was not compared at all, so protection stayed On while reads used the other layout. Both files record the layout their generation was encoded with, so a disagreement now fails the mount. * ec: reject an invalid recorded block size in weed fix -ecx A .vif with valid shard counts but a negative or unaligned block size was marked usable: a positive invalid value pinned the scan to a geometry that de-stripes to garbage, and a negative one ran the dual scan but left the invalid .vif in place afterwards. Validate it with the same rule the mount applies, and when it fails leave the layout unknown so the scan recovers it and the file is rewritten. * ec: validate the sidecar layout weed fix -ecx recovers from The .ecsum fallback was taken on DataShards > 0 alone, so a CRC-valid sidecar carrying the wrong generation, an incomplete ratio, or an unaligned block size would pin the reconstruction to one incorrect uniform-layout candidate instead of letting the dual scan decide. Require generation 0, a complete in-range ratio, and a valid block size; anything less leaves the layout unknown, which is the answer that still recovers by scanning. * ec: let only a genuinely absent sidecar choose the legacy layout With no .vif the bitrot sidecar is the only record of a volume's layout, and the mount fallback read a failed load, an unusable config, or a sidecar stamped for another generation as "assume legacy". A uniform generation-0 volume could therefore mount with legacy or another generation's geometry and answer reads with the wrong bytes. Present-but-unusable now fails the mount; only actual absence keeps the legacy defaults. Shared as EcShardConfigFromSidecar so every caller reads the sidecar the same way. * ec: treat a recorded-but-impossible layout as corruption, not as legacy - A .vif whose ecShardConfig is PRESENT but records an impossible ratio was answered with the default 10+4 and the legacy block layout, in both languages. That reads a uniform volume's shards at the wrong offsets and returns the wrong bytes. Only an entirely absent config still means "this predates the record"; a present one that cannot be true fails the mount. - The shard-count bound summed two uint32 counts as int, which wraps on a 32-bit build: 0x7fffffff + 0x7fffffff lands at -2 and slips under MaxShardCount. ValidEcShardCounts sums in uint64, and every EC call site that checked a recorded ratio now goes through it. * ec: rebuild on the geometry the sidecar records, and flag it when it disagrees The rebuild RPC passes BackgroundECContext, so RebuildEcFiles resolves the layout itself — and it resolved a missing or invalid .vif to the default 10+4 with the legacy block size. Two consequences: a 12+4 volume was reconstructed through a 10+4 matrix, which produces wrong bytes and never regenerates shards 14-15; and the chosen geometry then contradicted a valid uniform sidecar, which loadRebuildSidecar reported as BitrotOff — silently skipping the input and regenerated-shard checksum checks precisely when the volume had already lost its metadata. The layout now resolves from the bitrot sidecar (found across the server's disks, not just beside the base name) before falling back to the defaults, and a present-but-impossible ratio fails instead of being replaced. A sidecar that contradicts the chosen geometry is BitrotInvalid, which the existing unsafeIgnoreSidecar override still lets an operator push past. * ec: let the Rust rebuild read metadata off a sibling disk read_ec_shard_config searches only the location the rebuild writes into, so a volume whose .vif or generation-0 .ecsum sits on another of the server's disks resolved to the default 10+4 with the legacy block layout — the Rust half of the geometry-guessing the Go rebuild just stopped doing. It then reconstructs a custom-ratio or uniform volume through the wrong Reed-Solomon matrix and de-striping geometry. The rebuild now looks for the .vif in its own location and then each sibling, falls back to the generation-0 sidecar wherever that lives, and only defaults when neither exists anywhere. The encode-time .dat size the ecx rebuild needs is resolved the same way. * ec: resolve a rebuild's vif from every directory that may hold it RebuildEcFiles probed only <data-base>.vif. The caller knows the selected location's index directory and the sibling locations, but passed neither for metadata: additionalDirs carried shard directories only, and were searched for shards and the checksum sidecar. A split -dir/-dir.idx layout, or a disk holding only shards, therefore resolved a pre-sidecar custom-ratio volume to 10+4 and reconstructed through the wrong matrix — never regenerating shards 14-15. The caller now hands over the index and sibling directories, and the resolver probes the vif across all of them, matching what the Rust resolver already does for both the vif and the sidecar. * ec: make every rebuild consumer agree on the layout it resolved - The post-rebuild bitrot backfill re-derived the geometry from this directory's .vif alone and dropped the block size entirely, so a rebuild that resolved its layout from a sibling, the sidecar, or a uniform vif wrote a manifest describing a DIFFERENT layout — one later mounts reject, or that covers only the default shard count. The layout is resolved once now, through an exported ResolveRebuildECContext, and the rebuild and the backfill share that answer. - The Rust rebuild collected only each location's data directory, so a sibling's INDEX directory — where a split -dir/-dir.idx layout keeps .ecx/.ecj/.vif — was never probed, and a custom-ratio volume still resolved to 10+4 with the legacy layout. Both directories of every location are carried now, deduped against the rebuild's own. - A shard delivery can bring the checksum manifest with it, but the receive path only writes the file: a server that already had the volume mounted kept its resolved protection state (off) until a remount. The mount RPC re-resolves it once the shards it describes have been added. * ec: cover the rebuild's directory search with tests Reviewers flagged the sibling index directory twice, and the fix that closed it had no test of its own: the assembly sat inline in the rebuild handler, reachable only through a gRPC call against a populated store. Lifting it into rebuildSearchDirs / select_rebuild_location makes the rule assertable — a sibling contributes BOTH its data and its index directory, a shared index directory is listed once, and the rebuild's own data directory never repeats. Writing the Rust cases surfaced that the two implementations do not agree on where the rebuild's own index directory belongs, and both are right: Go's resolver takes a single directory list, so that directory has to be inside it, while Rust's takes the rebuild's data and index directories as their own arguments and would search them twice. The tests now state which contract each side is holding to, so neither drifts into the other's shape. Pure refactor otherwise; no behaviour change. * ec: search the index directory for the layout sidecar The Rust resolver looked for the generation-0 .ecsum in the rebuild's data directory and the sibling list, but not in the rebuild's own index directory — while the .vif lookup directly above it did, and Go's findBitrotSidecar has always checked both bases. On a split -dir/-dir.idx location that directory is where the metadata lives, and callers leave it out of the sibling list precisely because it is passed here separately, so nothing searched it. With no .vif anywhere the sidecar is the only surviving record of the layout. Missing it resolved a 12+4 uniform volume to 10+4 with the legacy striping — the test added here fails with (10, 4, 0) against the old code — and the rebuild then reconstructs through the wrong matrix and writes .ecx offsets that no reader can follow. * ec: let the rebuild see its own index directory The Rust rebuild takes a single flat directory list — the shape Go's RebuildEcFiles uses — so it cannot be handed the rebuild location's index directory separately the way the layout resolvers are, and the handler was passing the sibling list, which deliberately omits exactly that directory. On a split -dir/-dir.idx location that is where .ecx and .vif live, so the shard and index lookups could not see them. Go has always carried that directory in additionalDirs; this lines the two call sites up. * ec: let a config-free vif fall through to the layout sidecar A .vif that carries no ecShardConfig answers nothing about the layout, so it is no more informative than an absent one — but both trees treated its mere existence as the end of the search. Go went straight to the 10+4 legacy defaults without consulting the sidecar at all; Rust returned whatever ec_shard_config_from could make of a single directory. A 12+4 uniform volume with a legacy config-free vif therefore resolved as 10+4 legacy, and every read landed at the wrong shard offset. The sidecar lookup was also single-directory on both sides, while a split -dir/-dir.idx layout keeps .vif and .ecsum with the INDEX. Go's findBitrotSidecar has always taken both bases; the callers here passed only the data base, and the Rust bitrot resolver derived its path from the data base alone. Rust's layout resolver now takes a candidate directory list — data, index, then any siblings — and searches all of it, which also removes the early return that made the vif's presence decisive. load_vif_info_across_dirs reported `dir` even when load_vif_info had found the vif in `dir_idx`. Nothing reads that field today, so this changes no behaviour; it stops the next caller that resolves the rest of the volume's metadata against the answer from being sent to a disk holding none of it. Absence stays legal throughout: a volume with neither record is genuinely legacy. Present-but-unusable still fails the mount, now in the config-free-vif branch too. * ec: activate a delivered sidecar on every per-disk runtime A vid mounts as one EcVolume per disk, each with its own resolved protection state, but the post-delivery reload used the first-match lookup and so touched exactly one of them. The siblings kept reporting no protection until a remount — and since shard distribution deduplicates the metadata files onto the first target disk for a node, the runtime that got the .ecsum is not necessarily the one the lookup returns. Iterate every runtime instead, via a new FindAllEcVolumes and its Rust mut equivalent. Combined with each runtime now resolving its sidecar against its index directory as well as its data directory, a server sharing one -dir.idx across its disks activates all of them from the single delivered copy. The Rust volume server had no post-mount reload at all; it gets one here, matching Go. * ec: resolve the delivered sidecar across every EC metadata directory Reloading every per-disk runtime, added last round, did not by itself make the delivered manifest reachable. Startup mirroring copies .ecx/.ecj/.vif to every shard-bearing disk so each mounts self-contained, but deliberately not .ecsum, and a repair delivers exactly one copy. Each runtime was resolving against its own two directories, so every sibling of the disk that received the file kept reporting no protection however often it reloaded. Resolve one authoritative copy across every EC metadata directory instead of duplicating the file. Mirroring .ecsum would have to keep pace with a file that is rewritten as shards are repaired, and would not help the reported case at all: the delivery happens at runtime, and mirroring only runs at startup. The regression test pins both halves — a reload restricted to the volume's own directories still finds nothing, and the same reload given the server's metadata directories turns protection on. * ec: ask every directory before writing a TOFU baseline After a rebuild the opportunistic backfill asks whether this volume already has a checksum manifest, and answered from the data base alone. A split -dir/-dir.idx layout keeps the sidecar with the index, and a multi-disk server may keep it on a sibling, so an existing manifest read as absent. The consequence is worse than a missed read. On a false "no" the backfill writes a fresh sidecar at the data base from whatever the shards say right now — and the data base is the first candidate every resolver checks, so that TOFU baseline shadows the real manifest rather than sitting beside it. A shard that was silently corrupt gets blessed, and the record that would have caught it stops being consulted. FindBitrotSidecar exports the search the package already used internally, so the question is asked of the data base, the index base and the sibling disks — the same candidates the rebuild resolves its layout from. * ec: refuse a shard block size no encoder could have produced weed fix -ecx derived one from the raw shard extent, so a truncated or partially copied shard wrote a .vif that NewEcVolume then permanently refuses — the volume the tool was run to rescue could never mount again. An extent that is not a whole number of small blocks cannot have come from a uniform encode, so it is no longer offered as a candidate, and nothing unvalidated reaches the .vif. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: derive the .vif's dat size and block size from one measurement VolumeEcShardsGenerate stat'ed the .dat before the encode while WriteEcFiles stat'ed it again to size the blocks. A write landing between the two produced a .vif whose own two fields describe different files. WriteEcFiles now leaves both on the context, and fills a placeholder context in place so the caller can read them back. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: keep the source volume until every holder serves its shard layout The uniform layout rides in a .vif field older volume servers never knew: they discard it, mount the shards as legacy and return wrong bytes with nothing erroring, and the shard files are the same length either way so no other check notices. The upgrade order lived only in the release note. VolumeEcShardsInfo now reports the block size the holder actually serves, in both the Go and Rust servers, and the pre-delete verification refuses to drop the source unless every reachable holder echoes the one the shards were encoded with — while a rollback still exists. A server that predates the field answers 0, which is the negative answer. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: drop the rebuild's dead block-size parameters generateMissingEcFiles never reads largeBlockSize/smallBlockSize — Reed-Solomon reconstruction is layout-agnostic — so passing the legacy constants only advertised a layout the rebuild does not use. Also move UniformBlockSize's doc off ValidateBlockSize. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: warn about EC defaults only when the mount used them The "vif file not found, using defaults" warning fired even after the bitrot sidecar supplied a non-default layout, sending anyone triaging wrong bytes after the legacy layout the volume never mounted on. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: stat the distributed bitrot sidecar once The strict check re-stat'ed the file immediately before the stat that already gates inclusion, and a failed sidecar write now fails the encode outright, so the first could only fire on a deletion between the two lines. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: say what the reconstruct budget actually bounds A shard's buffer stays in bufs until its interval reconstructs, which is after the read that filled it released its permit, so the semaphore bounds round trips in flight and not retained bytes. Peak memory is the intervals reconstructing at once times the shards each reaches times the interval size. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * test: let the fake volume server report its delivered EC layout The pre-delete verification now asks each holder which shard block layout it serves, and a fake that always answered "unset" looked exactly like a volume server too old to know the field. Distribution ships the .vif to every holder alongside its shards, so read the layout back out of it as a real holder does. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7
993 lines
39 KiB
Go
993 lines
39 KiB
Go
package storage
|
|
|
|
import (
|
|
"context"
|
|
"errors"
|
|
"fmt"
|
|
"io"
|
|
"os"
|
|
"slices"
|
|
"sync"
|
|
"sync/atomic"
|
|
"time"
|
|
|
|
"github.com/klauspost/reedsolomon"
|
|
"golang.org/x/sync/semaphore"
|
|
|
|
"github.com/seaweedfs/seaweedfs/weed/glog"
|
|
"github.com/seaweedfs/seaweedfs/weed/operation"
|
|
"github.com/seaweedfs/seaweedfs/weed/pb"
|
|
"github.com/seaweedfs/seaweedfs/weed/pb/master_pb"
|
|
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
|
|
"github.com/seaweedfs/seaweedfs/weed/stats"
|
|
"github.com/seaweedfs/seaweedfs/weed/storage/erasure_coding"
|
|
"github.com/seaweedfs/seaweedfs/weed/storage/needle"
|
|
"github.com/seaweedfs/seaweedfs/weed/storage/types"
|
|
)
|
|
|
|
// errShardNotLocal indicates that the requested EC shard is simply not
|
|
// stored on this volume server. It is expected during normal reads when
|
|
// shards are spread across multiple servers, so callers should not log
|
|
// it as an error.
|
|
var errShardNotLocal = errors.New("ec shard not on this server")
|
|
|
|
// FindEcShardTargetLocation returns the disk that should receive a new
|
|
// shard / index file for (collection, vid). The selection order is:
|
|
//
|
|
// 1. a disk that already has the EC volume mounted (in-memory state),
|
|
// 2. a disk that owns the .ecx file on disk (volume not mounted yet),
|
|
// 3. any HDD with free space,
|
|
// 4. any disk with free space.
|
|
//
|
|
// Step 2 is the missing primitive that pinned subsequent shards to the
|
|
// first-shard disk during ec.rebuild. ec.rebuild only sets CopyEcxFile=true
|
|
// for the first shard, then relies on auto-select to land later shards on
|
|
// the same disk. Without an on-disk check, FindEcVolume returns nothing
|
|
// (no mount yet) and the fallback picks "any HDD with free space" — which
|
|
// can split shards from their index files across disks of the same node
|
|
// and lose them at startup. See issue #9212 and the orphan-shard
|
|
// reconciliation in #9244.
|
|
//
|
|
// dataShardCount is the data-shard count for this volume's EC layout (10
|
|
// for the OSS default, but custom ratios are supported via .vif). Callers
|
|
// pass it explicitly so this helper stays free of package-level constants
|
|
// — easier to mirror into builds that ship a different default ratio.
|
|
//
|
|
// Implementation walks s.Locations once and scores each disk by tier; the
|
|
// highest-tier disk wins, ties broken by free count. The earlier waterfall
|
|
// across four FindFreeLocation passes was equivalent but acquired
|
|
// volumesLock and ecVolumesLock RLocks (via VolumesLen / EcShardCount) up
|
|
// to four times per disk per call.
|
|
func (s *Store) FindEcShardTargetLocation(collection string, vid needle.VolumeId, dataShardCount int) *DiskLocation {
|
|
const (
|
|
tierAnyDisk = iota + 1
|
|
tierHDD
|
|
tierEcxOnDisk
|
|
tierMounted
|
|
)
|
|
|
|
var (
|
|
best *DiskLocation
|
|
bestTier int
|
|
bestFree int32
|
|
)
|
|
for _, loc := range s.Locations {
|
|
if loc.isDiskSpaceLow.Load() {
|
|
continue
|
|
}
|
|
freeCount := ecFreeShardCount(loc, dataShardCount)
|
|
if freeCount <= 0 {
|
|
continue
|
|
}
|
|
tier := tierAnyDisk
|
|
if loc.DiskType == types.HardDriveType {
|
|
tier = tierHDD
|
|
}
|
|
if loc.HasEcxFileOnDisk(collection, vid) {
|
|
tier = tierEcxOnDisk
|
|
}
|
|
if _, mounted := loc.FindEcVolume(vid); mounted {
|
|
tier = tierMounted
|
|
}
|
|
if best == nil || tier > bestTier || (tier == bestTier && freeCount > bestFree) {
|
|
best = loc
|
|
bestTier = tier
|
|
bestFree = freeCount
|
|
}
|
|
}
|
|
return best
|
|
}
|
|
|
|
// ecFreeShardCount returns the free EC shard capacity of loc, expressed
|
|
// in shard slots (not volume-equivalent slots). dataShardCount is the
|
|
// data-shard count of the EC layout being placed — see
|
|
// FindEcShardTargetLocation's docstring for why it's a parameter.
|
|
//
|
|
// FindFreeLocation in store.go does the same math but divides by
|
|
// DataShardsCount at the end. That truncation can exclude a disk that
|
|
// still has room for several individual shards (e.g. MaxVolumeCount=1,
|
|
// EcShardCount=1, dataShardCount=10 → reports 0 despite 9 free shard
|
|
// slots), which in this helper would re-route subsequent shards off the
|
|
// .ecx-owning disk and re-introduce the orphan-shard layout #9212 is
|
|
// trying to prevent. So we keep the result in shard slots throughout.
|
|
//
|
|
// MaxVolumeCount == 0 is the "unlimited" sentinel used elsewhere in the
|
|
// store (see hasFreeDiskLocation). Reporting a synthetic large free
|
|
// count keeps unlimited disks eligible while still letting tie-breaks
|
|
// prefer the less-loaded one.
|
|
func ecFreeShardCount(loc *DiskLocation, dataShardCount int) int32 {
|
|
if dataShardCount <= 0 {
|
|
return 0
|
|
}
|
|
if loc.MaxVolumeCount <= 0 {
|
|
const unlimitedFree = int32(1 << 30)
|
|
used := int32(loc.VolumesLen())*int32(dataShardCount) + int32(loc.EcShardCount())
|
|
if used >= unlimitedFree {
|
|
return 1
|
|
}
|
|
return unlimitedFree - used
|
|
}
|
|
free := (loc.MaxVolumeCount - int32(loc.VolumesLen())) * int32(dataShardCount)
|
|
free -= int32(loc.EcShardCount())
|
|
if free < 0 {
|
|
return 0
|
|
}
|
|
return free
|
|
}
|
|
|
|
func (s *Store) CollectErasureCodingHeartbeat() *master_pb.Heartbeat {
|
|
var ecShardMessages []*master_pb.VolumeEcShardInformationMessage
|
|
collectionEcShardSize := make(map[string]int64)
|
|
for diskId, location := range s.Locations {
|
|
location.ecVolumesLock.RLock()
|
|
for _, ecShards := range location.ecVolumes {
|
|
ecShardMessages = append(ecShardMessages, ecShards.ToVolumeEcShardInformationMessage(uint32(diskId))...)
|
|
|
|
for _, ecShard := range ecShards.Shards {
|
|
collectionEcShardSize[ecShards.Collection] += ecShard.Size()
|
|
}
|
|
}
|
|
location.ecVolumesLock.RUnlock()
|
|
}
|
|
|
|
for col, size := range collectionEcShardSize {
|
|
stats.VolumeServerDiskSizeGauge.WithLabelValues(col, "ec").Set(float64(size))
|
|
}
|
|
|
|
for col := range s.reportedEcCollections {
|
|
if _, stillHere := collectionEcShardSize[col]; !stillHere {
|
|
stats.VolumeServerDiskSizeGauge.DeleteLabelValues(col, "ec")
|
|
}
|
|
}
|
|
s.reportedEcCollections = make(map[string]struct{}, len(collectionEcShardSize))
|
|
for col := range collectionEcShardSize {
|
|
s.reportedEcCollections[col] = struct{}{}
|
|
}
|
|
|
|
return &master_pb.Heartbeat{
|
|
EcShards: ecShardMessages,
|
|
HasNoEcShards: len(ecShardMessages) == 0,
|
|
}
|
|
|
|
}
|
|
|
|
func (s *Store) MountEcShards(collection string, vid needle.VolumeId, shardId erasure_coding.ShardId, sourceDiskType string) error {
|
|
// The .ecx index file may live on a different disk than the one
|
|
// holding the .ec?? shard being mounted: ec.balance / ec.rebuild can
|
|
// place the .ecx on one local disk while later distributing shards
|
|
// across sibling disks of the same volume server. The per-disk
|
|
// IdxDirectory used by LoadEcShard would ENOENT the .ecx, so look up
|
|
// the .ecx owner across all DiskLocations once and route NewEcVolume
|
|
// at the directory that actually has the file. A 0-byte .ecx is
|
|
// treated as missing here (writeToFile can leave a stub on a failed
|
|
// EC distribute) so we still scan the rest of the disks.
|
|
ecxIdxDir, ecxFound := s.findEcxIdxDirForVolume(collection, vid)
|
|
|
|
// Collect failures so an all-disks-fail return reports every disk we
|
|
// tried rather than just the first one. Before this loop reordered
|
|
// itself to keep going after the first non-ENOENT error, a single
|
|
// shard-on-disk-without-.ecx situation would bail the loop and the
|
|
// operator saw "cannot open ec volume index" naming exactly one disk
|
|
// even when others held a valid index.
|
|
type diskError struct {
|
|
dir string
|
|
err error
|
|
}
|
|
var failures []diskError
|
|
|
|
for diskId, location := range s.Locations {
|
|
idxDir := location.IdxDirectory
|
|
if ecxFound {
|
|
// Fast path: if findEcxIdxDirForVolume already pointed at
|
|
// one of this disk's directories, the disk owns the .ecx
|
|
// and the local IdxDirectory is the right answer — skip
|
|
// the HasEcxFileOnDisk stat. Only fall back to the sibling
|
|
// disk's idxDir when this disk's directories are neither.
|
|
if location.IdxDirectory != ecxIdxDir && location.Directory != ecxIdxDir {
|
|
if !location.HasEcxFileOnDisk(collection, vid) {
|
|
idxDir = ecxIdxDir
|
|
}
|
|
}
|
|
}
|
|
ecVolume, err := location.loadEcShardWithIdxDir(collection, vid, shardId, idxDir)
|
|
if err == nil {
|
|
glog.V(0).Infof("MountEcShards %d.%d on disk ID %d", vid, shardId, diskId)
|
|
|
|
// Apply the orchestrator-supplied source disk type so the EC
|
|
// volume reports under it instead of the location's. Empty means
|
|
// "fall back to location's disk type" (#9423).
|
|
if sourceDiskType != "" {
|
|
ecVolume.SetDiskType(types.ToDiskType(sourceDiskType))
|
|
}
|
|
|
|
si := erasure_coding.NewShardsInfo()
|
|
si.Set(erasure_coding.NewShardInfo(shardId, erasure_coding.ShardSize(ecVolume.ShardSize())))
|
|
s.NewEcShardsChan <- &master_pb.VolumeEcShardInformationMessage{
|
|
Id: uint32(vid),
|
|
Collection: collection,
|
|
EcIndexBits: uint32(si.Bitmap()),
|
|
ShardSizes: si.SizesInt64(),
|
|
DiskType: string(ecVolume.DiskType()),
|
|
ExpireAtSec: ecVolume.ExpireAtSec,
|
|
DiskId: uint32(diskId),
|
|
EncodeTsNs: ecVolume.EncodeTsNs,
|
|
}
|
|
return nil
|
|
}
|
|
if errors.Is(err, os.ErrNotExist) {
|
|
// Shard or index not on this disk; another disk may own it.
|
|
continue
|
|
}
|
|
failures = append(failures, diskError{dir: location.Directory, err: err})
|
|
}
|
|
|
|
if len(failures) == 0 {
|
|
// No disk had the shard or the index; this volume server is not
|
|
// holding any artefacts for the requested shard. Name what we
|
|
// scanned for so the operator can tell "no .ecx anywhere" apart
|
|
// from "shard not on this server".
|
|
if !ecxFound {
|
|
return fmt.Errorf("MountEcShards %d.%d: no .ecx index found on any local disk", vid, shardId)
|
|
}
|
|
return fmt.Errorf("MountEcShards %d.%d not found on disk", vid, shardId)
|
|
}
|
|
// Some disks returned a real (non-ENOENT) error. Report them all so
|
|
// the caller can see whether the failures cluster around one disk
|
|
// (likely hardware) or are spread out (likely a config problem).
|
|
var b []byte
|
|
for i, f := range failures {
|
|
if i > 0 {
|
|
b = append(b, "; "...)
|
|
}
|
|
b = append(b, fmt.Sprintf("%s: %v", f.dir, f.err)...)
|
|
}
|
|
return fmt.Errorf("MountEcShards %d.%d load failures: %s", vid, shardId, string(b))
|
|
}
|
|
|
|
func (s *Store) UnmountEcShards(vid needle.VolumeId, shardId erasure_coding.ShardId, reqEncodeTsNs int64) error {
|
|
// Walk every disk: a split-disk reconciled volume can mount the same vid on
|
|
// more than one disk, so a first-match unmount would leave a sibling copy
|
|
// mounted and heartbeating. Emit one deletion delta per disk.
|
|
unmountedAny := false
|
|
var lastErr error
|
|
for diskId, location := range s.Locations {
|
|
ecShard, found := location.FindEcShard(vid, shardId)
|
|
if !found {
|
|
continue
|
|
}
|
|
// Capture the encode generation before unloading so the deletion delta
|
|
// carries it like the mount delta does.
|
|
var encodeTsNs int64
|
|
if ecVolume, ok := location.FindEcVolume(vid); ok {
|
|
encodeTsNs = ecVolume.EncodeTsNs
|
|
}
|
|
// Generation fence: when the caller carries a generation (stale-worker
|
|
// cleanup), only unmount a strictly-older generation; preserve a disk whose
|
|
// generation is same-or-newer, 0, or unknown, so a stale run cannot unmount
|
|
// a newer run's live shards. reqEncodeTsNs==0 (legacy/shell) unmounts all.
|
|
if reqEncodeTsNs > 0 && !(encodeTsNs > 0 && encodeTsNs < reqEncodeTsNs) {
|
|
glog.V(1).Infof("UnmountEcShards %d.%d disk_id:%d skipped: disk gen %d not older than request gen %d", vid, shardId, diskId, encodeTsNs, reqEncodeTsNs)
|
|
continue
|
|
}
|
|
if deleted := location.UnloadEcShard(vid, shardId); deleted {
|
|
si := erasure_coding.NewShardsInfo()
|
|
si.Set(erasure_coding.NewShardInfo(shardId, 0))
|
|
s.DeletedEcShardsChan <- &master_pb.VolumeEcShardInformationMessage{
|
|
Id: uint32(vid),
|
|
Collection: ecShard.Collection,
|
|
EcIndexBits: si.Bitmap(),
|
|
ShardSizes: si.SizesInt64(),
|
|
DiskType: string(ecShard.DiskType),
|
|
DiskId: uint32(diskId),
|
|
EncodeTsNs: encodeTsNs,
|
|
}
|
|
glog.V(0).Infof("UnmountEcShards %d.%d disk_id:%d", vid, shardId, diskId)
|
|
unmountedAny = true
|
|
} else {
|
|
lastErr = fmt.Errorf("UnmountEcShards %d.%d not found on disk %d", vid, shardId, diskId)
|
|
}
|
|
}
|
|
|
|
// nil when no disk held the shard (idempotent re-unmount).
|
|
if !unmountedAny {
|
|
return lastErr
|
|
}
|
|
return nil
|
|
}
|
|
|
|
func (s *Store) findEcShard(vid needle.VolumeId, shardId erasure_coding.ShardId) (diskId uint32, shard *erasure_coding.EcVolumeShard, found bool) {
|
|
for diskId, location := range s.Locations {
|
|
if v, found := location.FindEcShard(vid, shardId); found {
|
|
return uint32(diskId), v, found
|
|
}
|
|
}
|
|
return 0, nil, false
|
|
}
|
|
|
|
// FindEcShard returns the shard if any DiskLocation on this server holds it,
|
|
// along with that disk's id.
|
|
func (s *Store) FindEcShard(vid needle.VolumeId, shardId erasure_coding.ShardId) (diskId uint32, shard *erasure_coding.EcVolumeShard, found bool) {
|
|
return s.findEcShard(vid, shardId)
|
|
}
|
|
|
|
// FindEcVolumeWithShard returns the EcVolume on the disk that owns the given
|
|
// shard, plus the shard. The read guard must check the identity of the volume
|
|
// that owns the bytes served: on a multi-disk server one vid can hold shards
|
|
// from different encode runs across disks, so a first-match volume can differ.
|
|
func (s *Store) FindEcVolumeWithShard(vid needle.VolumeId, shardId erasure_coding.ShardId) (*erasure_coding.EcVolume, *erasure_coding.EcVolumeShard, bool) {
|
|
for _, location := range s.Locations {
|
|
if shard, found := location.FindEcShard(vid, shardId); found {
|
|
if ev, ok := location.FindEcVolume(vid); ok {
|
|
return ev, shard, true
|
|
}
|
|
}
|
|
}
|
|
return nil, nil, false
|
|
}
|
|
|
|
// EcMetadataDirs lists every directory on this server that could hold an EC
|
|
// volume's metadata — each disk's data and index directory. Startup mirroring
|
|
// gives each shard-bearing disk its own .ecx/.ecj/.vif, but the checksum
|
|
// sidecar is not mirrored and a repair delivers exactly one copy, so a runtime
|
|
// looking only at its own two directories cannot see it. Handing this list to
|
|
// the sidecar resolution keeps one authoritative copy reachable from every
|
|
// runtime rather than duplicating a file that is rewritten as shards are
|
|
// repaired and generations published.
|
|
func (s *Store) EcMetadataDirs() []string {
|
|
var dirs []string
|
|
appendDir := func(dir string) {
|
|
if dir == "" {
|
|
return
|
|
}
|
|
for _, existing := range dirs {
|
|
if existing == dir {
|
|
return
|
|
}
|
|
}
|
|
dirs = append(dirs, dir)
|
|
}
|
|
for _, location := range s.Locations {
|
|
appendDir(location.Directory)
|
|
appendDir(location.IdxDirectory)
|
|
}
|
|
return dirs
|
|
}
|
|
|
|
// FindAllEcVolumes returns every per-disk *EcVolume the store maps for vid. A vid
|
|
// can mount on N disks as N distinct runtimes, and the first-match FindEcVolume
|
|
// hides the siblings — so anything that has to reach the whole volume, rather
|
|
// than any one runtime of it, iterates this instead. Order mirrors the
|
|
// deterministic s.Locations order; nil when no disk holds the vid.
|
|
func (s *Store) FindAllEcVolumes(vid needle.VolumeId) []*erasure_coding.EcVolume {
|
|
var evs []*erasure_coding.EcVolume
|
|
for _, location := range s.Locations {
|
|
if ev, found := location.FindEcVolume(vid); found {
|
|
evs = append(evs, ev)
|
|
}
|
|
}
|
|
return evs
|
|
}
|
|
|
|
func (s *Store) FindEcVolume(vid needle.VolumeId) (*erasure_coding.EcVolume, bool) {
|
|
for _, location := range s.Locations {
|
|
if s, found := location.FindEcVolume(vid); found {
|
|
return s, true
|
|
}
|
|
}
|
|
return nil, false
|
|
}
|
|
|
|
// FindEcVolumeDiskIds returns every disk_id on this store that has an
|
|
// EcVolume entry for the given volume. Useful for diagnostic logging
|
|
// when a single FindEcVolume hit hides which disk is actually holding
|
|
// the mount (e.g., the ReceiveFile mounted-volume guard).
|
|
func (s *Store) FindEcVolumeDiskIds(vid needle.VolumeId) []uint32 {
|
|
var ids []uint32
|
|
for diskId, location := range s.Locations {
|
|
if _, found := location.FindEcVolume(vid); found {
|
|
ids = append(ids, uint32(diskId))
|
|
}
|
|
}
|
|
return ids
|
|
}
|
|
|
|
// shardFiles is a list of shard files, which is used to return the shard locations
|
|
func (s *Store) CollectEcShards(vid needle.VolumeId, shardFileNames []string) (ecVolume *erasure_coding.EcVolume, found bool) {
|
|
for _, location := range s.Locations {
|
|
if s, foundShards := location.CollectEcShards(vid, shardFileNames); foundShards {
|
|
ecVolume = s
|
|
found = true
|
|
}
|
|
}
|
|
return
|
|
}
|
|
|
|
func (s *Store) DestroyEcVolume(vid needle.VolumeId) {
|
|
for _, location := range s.Locations {
|
|
location.DestroyEcVolume(vid)
|
|
}
|
|
}
|
|
|
|
// UnloadEcVolume drops any in-memory EcVolume for vid from every disk and closes
|
|
// its fds without deleting files, so a following unlink frees the inodes.
|
|
func (s *Store) UnloadEcVolume(vid needle.VolumeId) {
|
|
for _, location := range s.Locations {
|
|
location.unloadEcVolume(vid)
|
|
}
|
|
}
|
|
|
|
func (s *Store) ReadEcShardNeedle(vid needle.VolumeId, n *needle.Needle, onReadSizeFn func(size types.Size)) (int, error) {
|
|
for _, location := range s.Locations {
|
|
if localEcVolume, found := location.FindEcVolume(vid); found {
|
|
|
|
offset, size, intervals, err := localEcVolume.LocateEcShardNeedle(n.Id, localEcVolume.Version)
|
|
if err != nil {
|
|
return 0, fmt.Errorf("locate in local ec volume: %w", err)
|
|
}
|
|
if size.IsDeleted() {
|
|
return 0, ErrorDeleted
|
|
}
|
|
|
|
if onReadSizeFn != nil {
|
|
onReadSizeFn(size)
|
|
}
|
|
|
|
glog.V(3).Infof("read ec volume %d offset %d size %d intervals:%+v", vid, offset.ToActualOffset(), size, intervals)
|
|
|
|
if len(intervals) > 1 {
|
|
glog.V(3).Infof("ReadEcShardNeedle needle id %s intervals:%+v", n.String(), intervals)
|
|
}
|
|
bytes, isDeleted, err := s.readEcShardIntervals(n.Id, localEcVolume, intervals)
|
|
// A holder reporting the needle deleted is authoritative -- deletes are
|
|
// never invented and never undone -- so answer that ahead of whatever
|
|
// error the shards it could not gather produced.
|
|
if isDeleted {
|
|
return 0, ErrorDeleted
|
|
}
|
|
if err != nil {
|
|
return 0, fmt.Errorf("ReadEcShardIntervals: %w", err)
|
|
}
|
|
|
|
err = n.ReadBytes(bytes, offset.ToActualOffset(), size, localEcVolume.Version)
|
|
if err != nil {
|
|
return 0, fmt.Errorf("ec volume %d needle %s offset %d size %d: %w", vid, n.String(), offset.ToActualOffset(), size, err)
|
|
}
|
|
|
|
return len(bytes), nil
|
|
}
|
|
}
|
|
return 0, fmt.Errorf("ec shard %d not found", vid)
|
|
}
|
|
|
|
var (
|
|
// ecRecoverBudget bounds the interval-sized buffers EC recovery holds across
|
|
// all concurrent reads: a peer that is slow to fail keeps a whole fan-out of
|
|
// them alive for the gRPC timeout, and a burst of those walked servers into an
|
|
// OOM. A var so a test can narrow it alongside ecRecoverSem.
|
|
ecRecoverBudget int64 = 256 << 20
|
|
ecRecoverSem = semaphore.NewWeighted(ecRecoverBudget)
|
|
)
|
|
|
|
// ecIntervalReadConcurrency bounds the fan-out of a single needle read. Blocks
|
|
// that follow each other in the .dat live on different shards, so a needle
|
|
// spanning several of them costs one round trip per block when read in sequence.
|
|
const ecIntervalReadConcurrency = 8
|
|
|
|
func (s *Store) readEcShardIntervals(needleId types.NeedleId, ecVolume *erasure_coding.EcVolume, intervals []erasure_coding.Interval) (data []byte, is_deleted bool, err error) {
|
|
if err = s.cachedLookupEcShardLocations(ecVolume); err != nil {
|
|
return nil, false, fmt.Errorf("failed to locate shard via master grpc %s: %v", s.MasterAddress, err)
|
|
}
|
|
|
|
var totalSize int
|
|
for _, interval := range intervals {
|
|
totalSize += int(interval.Size)
|
|
}
|
|
data = make([]byte, totalSize)
|
|
|
|
if len(intervals) <= 1 {
|
|
for _, interval := range intervals {
|
|
if is_deleted, err = s.readOneEcShardInterval(needleId, ecVolume, interval, data); err != nil {
|
|
return nil, is_deleted, err
|
|
}
|
|
}
|
|
return data, is_deleted, nil
|
|
}
|
|
|
|
errs := make([]error, len(intervals))
|
|
var deleted atomic.Bool
|
|
var wg sync.WaitGroup
|
|
sem := make(chan struct{}, ecIntervalReadConcurrency)
|
|
var pos int
|
|
for i, interval := range intervals {
|
|
buf := data[pos : pos+int(interval.Size)]
|
|
pos += int(interval.Size)
|
|
wg.Add(1)
|
|
sem <- struct{}{}
|
|
go func() {
|
|
defer wg.Done()
|
|
defer func() { <-sem }()
|
|
isDeleted, e := s.readOneEcShardInterval(needleId, ecVolume, interval, buf)
|
|
if isDeleted {
|
|
deleted.Store(true)
|
|
}
|
|
errs[i] = e
|
|
}()
|
|
}
|
|
wg.Wait()
|
|
|
|
is_deleted = deleted.Load()
|
|
for _, e := range errs {
|
|
if e != nil {
|
|
return nil, is_deleted, e
|
|
}
|
|
}
|
|
return data, is_deleted, nil
|
|
}
|
|
|
|
// readOneEcShardInterval fills data, which must be interval.Size long.
|
|
func (s *Store) readOneEcShardInterval(needleId types.NeedleId, ecVolume *erasure_coding.EcVolume, interval erasure_coding.Interval, data []byte) (is_deleted bool, err error) {
|
|
shardId, actualOffset := ecVolume.IntervalToShardIdAndOffset(interval)
|
|
|
|
// try local read
|
|
err = s.readLocalEcShardInterval(ecVolume, shardId, data, actualOffset)
|
|
if err == nil {
|
|
return
|
|
}
|
|
if errors.Is(err, errShardNotLocal) {
|
|
// expected when shards are spread across servers; fall through to remote read
|
|
glog.V(4).Infof("ec shard %d.%d not local, will try remote", ecVolume.VolumeId, shardId)
|
|
} else {
|
|
glog.V(0).Infof("read local ec shard %d.%d offset %d: %v", ecVolume.VolumeId, shardId, actualOffset, err)
|
|
}
|
|
|
|
ecVolume.ShardLocationsLock.RLock()
|
|
sourceDataNodes, hasShardIdLocation := ecVolume.ShardLocations[shardId]
|
|
ecVolume.ShardLocationsLock.RUnlock()
|
|
|
|
// try reading directly
|
|
if hasShardIdLocation {
|
|
_, is_deleted, err = s.readRemoteEcShardInterval(sourceDataNodes, needleId, ecVolume.VolumeId, shardId, data, actualOffset, ecVolume.EncodeTsNs)
|
|
if err == nil {
|
|
return
|
|
}
|
|
glog.V(0).Infof("read remote ec shard %d.%d locations: %v", ecVolume.VolumeId, shardId, err)
|
|
// Recovery below skips this very shard, so nothing else invalidates the
|
|
// location that just failed -- and a shard that has moved would otherwise
|
|
// be reconstructed on every read until the map's own window expires.
|
|
markShardLocationsStale(ecVolume)
|
|
}
|
|
|
|
// try reading by recovering from other shards
|
|
_, is_deleted, err = s.recoverOneRemoteEcShardInterval(needleId, ecVolume, shardId, data, actualOffset)
|
|
if err == nil {
|
|
return
|
|
}
|
|
glog.V(0).Infof("recover ec shard %d.%d : %v", ecVolume.VolumeId, shardId, err)
|
|
|
|
return
|
|
}
|
|
|
|
func forgetShardId(ecVolume *erasure_coding.EcVolume, shardId erasure_coding.ShardId) {
|
|
// failed to access the source data nodes, clear it up
|
|
ecVolume.ShardLocationsLock.Lock()
|
|
delete(ecVolume.ShardLocations, shardId)
|
|
ecVolume.ShardLocationsStale = true
|
|
ecVolume.ShardLocationsLock.Unlock()
|
|
}
|
|
|
|
// markShardLocationsStale flags the cached map for a prompt re-check after a
|
|
// read failed against one of its locations. Unlike forgetShardId it keeps the
|
|
// entry: a direct read is worth retrying, since a dead peer fails fast.
|
|
func markShardLocationsStale(ecVolume *erasure_coding.EcVolume) {
|
|
ecVolume.ShardLocationsLock.Lock()
|
|
ecVolume.ShardLocationsStale = true
|
|
ecVolume.ShardLocationsLock.Unlock()
|
|
}
|
|
|
|
// ecShardLocationsTTL is how long a cached shard map is trusted. A complete map
|
|
// is trusted longest. One short of DataShards, or one a failed read has just
|
|
// invalidated, is re-checked promptly: until it is, every read of the dropped
|
|
// shard skips the direct fetch and pays for a Reed-Solomon recovery instead.
|
|
func ecShardLocationsTTL(shardCount int, stale bool, ecCtx *erasure_coding.ECContext) time.Duration {
|
|
switch {
|
|
case stale || shardCount < ecCtx.DataShards:
|
|
return 11 * time.Second
|
|
case shardCount == ecCtx.Total():
|
|
return 37 * time.Minute
|
|
default:
|
|
return 7 * time.Minute
|
|
}
|
|
}
|
|
|
|
func (s *Store) cachedLookupEcShardLocations(ecVolume *erasure_coding.EcVolume) (err error) {
|
|
|
|
// Use the volume's own EC ratio so a custom-ratio volume (e.g. 9+3) is judged
|
|
// complete/recoverable against its real data-shard count, not the build default.
|
|
// In OSS the ratio is always 10+4, so this is a no-op.
|
|
ecCtx := ecVolume.ECContext
|
|
if ecCtx == nil {
|
|
ecCtx = erasure_coding.NewDefaultECContext(ecVolume.Collection, ecVolume.VolumeId)
|
|
}
|
|
|
|
// Judge the map and consume its mark in one critical section, so a mark raised
|
|
// from here on belongs to the next refresh rather than being cleared by this
|
|
// one -- the read that raised it has disproved the map this lookup is about to
|
|
// install. Recover goroutines mutate all three fields via forgetShardId, so an
|
|
// unguarded read would race a concurrent map write besides.
|
|
ecVolume.ShardLocationsLock.Lock()
|
|
stale := ecVolume.ShardLocationsStale
|
|
ttl := ecShardLocationsTTL(len(ecVolume.ShardLocations), stale, ecCtx)
|
|
fresh := ecVolume.ShardLocationsRefreshTime.Add(ttl).After(time.Now())
|
|
if !fresh {
|
|
ecVolume.ShardLocationsStale = false
|
|
}
|
|
ecVolume.ShardLocationsLock.Unlock()
|
|
if fresh {
|
|
return nil
|
|
}
|
|
|
|
glog.V(3).Infof("lookup and cache ec volume %d locations", ecVolume.VolumeId)
|
|
|
|
err = operation.WithMasterServerClient(context.Background(), false, s.MasterAddress, s.grpcDialOption, func(masterClient master_pb.SeaweedClient) error {
|
|
req := &master_pb.LookupEcVolumeRequest{
|
|
VolumeId: uint32(ecVolume.VolumeId),
|
|
}
|
|
resp, err := masterClient.LookupEcVolume(context.Background(), req)
|
|
if err != nil {
|
|
return fmt.Errorf("lookup ec volume %d: %v", ecVolume.VolumeId, err)
|
|
}
|
|
if len(resp.ShardIdLocations) < ecCtx.DataShards {
|
|
return fmt.Errorf("only %d shards found but %d required", len(resp.ShardIdLocations), ecCtx.DataShards)
|
|
}
|
|
|
|
ecVolume.ShardLocationsLock.Lock()
|
|
for _, shardIdLocations := range resp.ShardIdLocations {
|
|
shardId := erasure_coding.ShardId(shardIdLocations.ShardId)
|
|
delete(ecVolume.ShardLocations, shardId)
|
|
for _, loc := range shardIdLocations.Locations {
|
|
ecVolume.ShardLocations[shardId] = append(ecVolume.ShardLocations[shardId], pb.NewServerAddressFromLocation(loc))
|
|
}
|
|
}
|
|
ecVolume.ShardLocationsRefreshTime = time.Now()
|
|
ecVolume.ShardLocationsLock.Unlock()
|
|
|
|
return nil
|
|
})
|
|
if err != nil {
|
|
// The lookup this mark was consumed for did not answer, and the refresh time
|
|
// is only advanced on success -- so without putting the mark back, a map a
|
|
// read had already disproved would be trusted for its full window again.
|
|
// Marking unconditionally is safe: where no mark was consumed the refresh
|
|
// time is stale anyway, and the next call looks up whatever the mark says.
|
|
markShardLocationsStale(ecVolume)
|
|
}
|
|
return
|
|
}
|
|
|
|
func (s *Store) readLocalEcShardInterval(ecVolume *erasure_coding.EcVolume, shardId erasure_coding.ShardId, buf []byte, offset int64) error {
|
|
// Resolve the shard together with the EcVolume on the disk that owns it; the
|
|
// shard may live on a sibling disk of this server.
|
|
ownerVolume, shard, found := s.FindEcVolumeWithShard(ecVolume.VolumeId, shardId)
|
|
if !found {
|
|
return fmt.Errorf("shard %d for volume %d: %w", shardId, ecVolume.VolumeId, errShardNotLocal)
|
|
}
|
|
// Skip a local shard whose identity doesn't match the caller's index, so the
|
|
// read recovers from the correct generation. Lenient only when the caller has
|
|
// no identity (pre-upgrade): a known caller must not accept an unstamped local
|
|
// shard, which would serve a stale pre-upgrade generation.
|
|
if ecVolume.EncodeTsNs != 0 && ecVolume.EncodeTsNs != ownerVolume.EncodeTsNs {
|
|
glog.V(1).Infof("skip local ec shard %d.%d from a different encode run: caller EncodeTsNs %d, local %d", ecVolume.VolumeId, shardId, ecVolume.EncodeTsNs, ownerVolume.EncodeTsNs)
|
|
return fmt.Errorf("shard %d for volume %d: %w", shardId, ecVolume.VolumeId, errShardNotLocal)
|
|
}
|
|
|
|
readBytes, err := shard.ReadAt(buf, offset)
|
|
if err != nil {
|
|
return fmt.Errorf("failed to read local EC shard %d for volume %d: %v", shardId, ecVolume.VolumeId, err)
|
|
}
|
|
if got, want := readBytes, len(buf); got != want {
|
|
return fmt.Errorf("expected %d bytes for local EC shard %d on volume %d, got %d", want, shardId, ecVolume.VolumeId, got)
|
|
}
|
|
|
|
return nil
|
|
}
|
|
|
|
func (s *Store) readRemoteEcShardInterval(sourceDataNodes []pb.ServerAddress, needleId types.NeedleId, vid needle.VolumeId, shardId erasure_coding.ShardId, buf []byte, offset int64, expectedEncodeTsNs int64) (n int, is_deleted bool, err error) {
|
|
|
|
if len(sourceDataNodes) == 0 {
|
|
return 0, false, fmt.Errorf("failed to find ec shard %d.%d", vid, shardId)
|
|
}
|
|
|
|
for _, sourceDataNode := range sourceDataNodes {
|
|
glog.V(3).Infof("read remote ec shard %d.%d from %s", vid, shardId, sourceDataNode)
|
|
n, is_deleted, err = s.doReadRemoteEcShardInterval(sourceDataNode, needleId, vid, shardId, buf, offset, expectedEncodeTsNs)
|
|
if err == nil {
|
|
return
|
|
}
|
|
glog.V(1).Infof("read remote ec shard %d.%d from %s: %v", vid, shardId, sourceDataNode, err)
|
|
}
|
|
|
|
return
|
|
}
|
|
|
|
func (s *Store) doReadRemoteEcShardInterval(sourceDataNode pb.ServerAddress, needleId types.NeedleId, vid needle.VolumeId, shardId erasure_coding.ShardId, buf []byte, offset int64, expectedEncodeTsNs int64) (n int, is_deleted bool, err error) {
|
|
|
|
err = operation.WithVolumeServerClient(false, sourceDataNode, s.grpcDialOption, func(client volume_server_pb.VolumeServerClient) error {
|
|
|
|
// copy data slice
|
|
shardReadClient, err := client.VolumeEcShardRead(context.Background(), &volume_server_pb.VolumeEcShardReadRequest{
|
|
VolumeId: uint32(vid),
|
|
ShardId: uint32(shardId),
|
|
Offset: offset,
|
|
Size: int64(len(buf)),
|
|
FileKey: uint64(needleId),
|
|
EncodeTsNs: expectedEncodeTsNs,
|
|
})
|
|
if err != nil {
|
|
return fmt.Errorf("failed to start reading ec shard %d.%d from %s: %v", vid, shardId, sourceDataNode, err)
|
|
}
|
|
|
|
for {
|
|
resp, receiveErr := shardReadClient.Recv()
|
|
if receiveErr == io.EOF {
|
|
break
|
|
}
|
|
if receiveErr != nil {
|
|
return fmt.Errorf("receiving ec shard %d.%d from %s: %v", vid, shardId, sourceDataNode, receiveErr)
|
|
}
|
|
// Validate the served shard's identity client-side, so the guard holds
|
|
// even against a pre-upgrade server that ignored the request field (it
|
|
// returns 0). A mismatch fails the read; the caller recovers from parity.
|
|
if expectedEncodeTsNs != 0 && resp.EncodeTsNs != expectedEncodeTsNs {
|
|
return fmt.Errorf("ec shard %d.%d from %s belongs to a different encode run (want %d, got %d)", vid, shardId, sourceDataNode, expectedEncodeTsNs, resp.EncodeTsNs)
|
|
}
|
|
if resp.IsDeleted {
|
|
is_deleted = true
|
|
}
|
|
copy(buf[n:n+len(resp.Data)], resp.Data)
|
|
n += len(resp.Data)
|
|
}
|
|
|
|
return nil
|
|
})
|
|
if err != nil {
|
|
return 0, is_deleted, fmt.Errorf("read ec shard %d.%d from %s: %v", vid, shardId, sourceDataNode, err)
|
|
}
|
|
|
|
// A non-deleted interval must arrive whole: the server stamps EncodeTsNs only
|
|
// on chunks that carry bytes, so a short or empty stream (e.g. immediate EOF
|
|
// from a pre-upgrade or stale server) leaves the buffer partly zero-filled and
|
|
// unvalidated. Reject it so the caller recovers from parity. The is_deleted
|
|
// short-circuit legitimately returns n=0 with no data and is exempt, matching
|
|
// readLocalEcShardInterval's got==len(buf) rule for the local path.
|
|
if !is_deleted && n != len(buf) {
|
|
return n, is_deleted, fmt.Errorf("short read ec shard %d.%d from %s: got %d want %d", vid, shardId, sourceDataNode, n, len(buf))
|
|
}
|
|
|
|
return
|
|
}
|
|
|
|
// gatherEcShardIntervals collects the same interval from every shard but shardIdToRecover,
|
|
// returning one buffer per shard and nil where the shard could not be gathered. It stops
|
|
// once DataShards of them are in hand: reconstruction consumes no more than that.
|
|
func (s *Store) gatherEcShardIntervals(needleId types.NeedleId, ecVolume *erasure_coding.EcVolume, ecCtx *erasure_coding.ECContext, shardIdToRecover erasure_coding.ShardId, size int, offset int64) (shardIntervals [][]byte, isDeleted bool) {
|
|
shardIntervals = make([][]byte, ecCtx.Total())
|
|
|
|
// A shard this server already holds costs no round trip and no peer buffer,
|
|
// so seed those before asking peers for the rest.
|
|
available := 0
|
|
for shardId := erasure_coding.ShardId(0); int(shardId) < ecCtx.Total() && available < ecCtx.DataShards; shardId++ {
|
|
if shardId == shardIdToRecover {
|
|
continue
|
|
}
|
|
if _, _, found := s.FindEcVolumeWithShard(ecVolume.VolumeId, shardId); !found {
|
|
continue
|
|
}
|
|
data := make([]byte, size)
|
|
if localErr := s.readLocalEcShardInterval(ecVolume, shardId, data, offset); localErr != nil {
|
|
glog.V(3).Infof("recover: read local ec shard %d.%d: %v", ecVolume.VolumeId, shardId, localErr)
|
|
continue
|
|
}
|
|
shardIntervals[shardId] = data
|
|
available++
|
|
}
|
|
|
|
var candidates []erasure_coding.ShardId
|
|
candidateLocations := make(map[erasure_coding.ShardId][]pb.ServerAddress)
|
|
ecVolume.ShardLocationsLock.RLock()
|
|
for shardId, locations := range ecVolume.ShardLocations {
|
|
|
|
// skip the shard being recovered, one already seeded locally, or an empty shard
|
|
if shardId == shardIdToRecover || int(shardId) >= ecCtx.Total() || shardIntervals[shardId] != nil {
|
|
continue
|
|
}
|
|
if len(locations) == 0 {
|
|
glog.V(3).Infof("readRemoteEcShardInterval missing %d.%d from %+v", ecVolume.VolumeId, shardId, locations)
|
|
continue
|
|
}
|
|
candidates = append(candidates, shardId)
|
|
candidateLocations[shardId] = locations
|
|
}
|
|
ecVolume.ShardLocationsLock.RUnlock()
|
|
|
|
// The recover goroutines run concurrently, so the deleted flag is collected
|
|
// atomically and folded into the named return after they join, rather than each
|
|
// goroutine writing the shared bool directly.
|
|
var isDeletedFlag atomic.Bool
|
|
|
|
// Reconstruction consumes DataShards buffers, so reading every remaining shard
|
|
// holds a third more than that and asks a third more of peers that are, by
|
|
// then, already struggling. Widen only when a wave falls short.
|
|
for len(candidates) > 0 && available < ecCtx.DataShards {
|
|
wave := min(ecCtx.DataShards-available, len(candidates))
|
|
var wg sync.WaitGroup
|
|
var fetched atomic.Int64
|
|
for _, shardId := range candidates[:wave] {
|
|
locations := candidateLocations[shardId]
|
|
wg.Add(1)
|
|
go func() {
|
|
defer wg.Done()
|
|
data := make([]byte, size)
|
|
nRead, isDeleted, readErr := s.readRemoteEcShardInterval(locations, needleId, ecVolume.VolumeId, shardId, data, offset, ecVolume.EncodeTsNs)
|
|
if readErr != nil {
|
|
glog.V(3).Infof("recover: readRemoteEcShardInterval %d.%d %d bytes from %+v: %v", ecVolume.VolumeId, shardId, nRead, locations, readErr)
|
|
forgetShardId(ecVolume, shardId)
|
|
}
|
|
if isDeleted {
|
|
isDeletedFlag.Store(true)
|
|
}
|
|
if nRead == size {
|
|
shardIntervals[shardId] = data
|
|
fetched.Add(1)
|
|
}
|
|
}()
|
|
}
|
|
wg.Wait()
|
|
candidates = candidates[wave:]
|
|
available += int(fetched.Load())
|
|
if isDeletedFlag.Load() {
|
|
// every shard of a deleted needle answers deleted, so another wave cannot help
|
|
break
|
|
}
|
|
}
|
|
|
|
return shardIntervals, isDeletedFlag.Load()
|
|
}
|
|
|
|
// checkEcShardRebuildable rejects a parity target. ReconstructData rebuilds data shards
|
|
// only, so it leaves a parity slot nil and reports no error, and the caller would copy
|
|
// that out as a zero-filled buffer and report a successful read.
|
|
func checkEcShardRebuildable(ecVolume *erasure_coding.EcVolume, ecCtx *erasure_coding.ECContext, shardIdToRecover erasure_coding.ShardId) error {
|
|
if int(shardIdToRecover) >= ecCtx.DataShards {
|
|
return fmt.Errorf("cannot reconstruct shard %d.%d: only data shards can be rebuilt, %d of %d are parity",
|
|
ecVolume.VolumeId, shardIdToRecover, ecCtx.ParityShards, ecCtx.Total())
|
|
}
|
|
return nil
|
|
}
|
|
|
|
// reconstructEcShardInterval rebuilds one shard's interval in place from the others.
|
|
func reconstructEcShardInterval(ecVolume *erasure_coding.EcVolume, ecCtx *erasure_coding.ECContext, shardIntervals [][]byte, shardIdToRecover erasure_coding.ShardId) error {
|
|
if err := checkEcShardRebuildable(ecVolume, ecCtx, shardIdToRecover); err != nil {
|
|
return err
|
|
}
|
|
|
|
enc, err := reedsolomon.New(ecCtx.DataShards, ecCtx.ParityShards)
|
|
if err != nil {
|
|
return fmt.Errorf("failed to create encoder: %w", err)
|
|
}
|
|
|
|
// Count and log available shards for diagnostics
|
|
availableShards := make([]erasure_coding.ShardId, 0, ecCtx.Total())
|
|
missingShards := make([]erasure_coding.ShardId, 0, ecCtx.ParityShards+1)
|
|
for shardId := 0; shardId < ecCtx.Total(); shardId++ {
|
|
if shardIntervals[shardId] != nil {
|
|
availableShards = append(availableShards, erasure_coding.ShardId(shardId))
|
|
} else {
|
|
missingShards = append(missingShards, erasure_coding.ShardId(shardId))
|
|
}
|
|
}
|
|
|
|
glog.V(3).Infof("recover ec shard %d.%d: %d shards available %v, %d missing %v",
|
|
ecVolume.VolumeId, shardIdToRecover,
|
|
len(availableShards), availableShards,
|
|
len(missingShards), missingShards)
|
|
|
|
if len(availableShards) < ecCtx.DataShards {
|
|
return fmt.Errorf("cannot recover shard %d.%d: only %d shards available %v, need at least %d (missing: %v)",
|
|
ecVolume.VolumeId, shardIdToRecover,
|
|
len(availableShards), availableShards,
|
|
ecCtx.DataShards, missingShards)
|
|
}
|
|
|
|
// Rebuild only what was asked for. ReconstructData rebuilds every missing data
|
|
// shard, and a gather that stopped at DataShards can leave up to ParityShards of
|
|
// them missing -- an interval-sized buffer and a decode each, discarded unread.
|
|
// The mask is Total() long, not DataShards: reedsolomon documents both lengths
|
|
// but indexes the short one past its end when a parity shard is absent, which
|
|
// here it usually is.
|
|
required := make([]bool, ecCtx.Total())
|
|
required[shardIdToRecover] = true
|
|
if err := enc.ReconstructSome(shardIntervals, required); err != nil {
|
|
return fmt.Errorf("failed to reconstruct data for shard %d.%d with %d available shards %v: %w",
|
|
ecVolume.VolumeId, shardIdToRecover, len(availableShards), availableShards, err)
|
|
}
|
|
|
|
return nil
|
|
}
|
|
|
|
func (s *Store) recoverOneRemoteEcShardInterval(needleId types.NeedleId, ecVolume *erasure_coding.EcVolume, shardIdToRecover erasure_coding.ShardId, buf []byte, offset int64) (n int, is_deleted bool, err error) {
|
|
glog.V(3).Infof("recover ec shard %d.%d from other locations", ecVolume.VolumeId, shardIdToRecover)
|
|
|
|
// Reconstruct with the volume's OWN EC ratio (loaded from its .vif), not the
|
|
// build default, so a custom-ratio volume (e.g. 9+3) is decoded with the matrix
|
|
// that actually produced its shards -- decoding it as 10+4 would corrupt the
|
|
// recovered bytes. In OSS the ratio is always 10+4, so this is a no-op.
|
|
ecCtx := ecVolume.ECContext
|
|
if ecCtx == nil {
|
|
ecCtx = erasure_coding.NewDefaultECContext(ecVolume.Collection, ecVolume.VolumeId)
|
|
}
|
|
// checked before the gather: a doomed target should not cost a fan-out, nor drop a
|
|
// sibling's location through forgetShardId on the way to failing
|
|
if err := checkEcShardRebuildable(ecVolume, ecCtx, shardIdToRecover); err != nil {
|
|
return 0, false, err
|
|
}
|
|
|
|
// Charge the buffers this recovery is about to hold against the budget, so a
|
|
// burst of them queues here rather than on the heap: DataShards gathered, plus
|
|
// the one the rebuild allocates for the shard it recreates.
|
|
weight := int64(len(buf)) * int64(ecCtx.DataShards+1)
|
|
if weight > ecRecoverBudget {
|
|
// An interval whose fan-out outgrows the whole budget takes all of it and
|
|
// so runs alone, rather than blocking forever on an acquire that can never
|
|
// succeed. The cap is then one such recovery, not a burst of them.
|
|
weight = ecRecoverBudget
|
|
}
|
|
if err = ecRecoverSem.Acquire(context.Background(), weight); err != nil {
|
|
return 0, false, err
|
|
}
|
|
defer ecRecoverSem.Release(weight)
|
|
|
|
shardIntervals, is_deleted := s.gatherEcShardIntervals(needleId, ecVolume, ecCtx, shardIdToRecover, len(buf), offset)
|
|
if err := reconstructEcShardInterval(ecVolume, ecCtx, shardIntervals, shardIdToRecover); err != nil {
|
|
return 0, is_deleted, err
|
|
}
|
|
glog.V(4).Infof("recovered ec shard %d.%d from other locations", ecVolume.VolumeId, shardIdToRecover)
|
|
|
|
copy(buf, shardIntervals[shardIdToRecover])
|
|
|
|
return len(buf), is_deleted, nil
|
|
}
|
|
|
|
func (s *Store) EcVolumes() (ecVolumes []*erasure_coding.EcVolume) {
|
|
for _, location := range s.Locations {
|
|
location.ecVolumesLock.RLock()
|
|
for _, v := range location.ecVolumes {
|
|
ecVolumes = append(ecVolumes, v)
|
|
}
|
|
location.ecVolumesLock.RUnlock()
|
|
}
|
|
slices.SortFunc(ecVolumes, func(a, b *erasure_coding.EcVolume) int {
|
|
return int(a.VolumeId) - int(b.VolumeId)
|
|
})
|
|
return ecVolumes
|
|
}
|