mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-08 23:37:43 +02:00
1df165d514ae9ef8b278ea340582af70e7ca8bdf
15415
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1df165d514 |
fix(volume): keep the TTL clock across a vacuum commit instead of rescanning (#11630)
* fix(volume): keep the TTL clock across a vacuum commit instead of rescanning CommitCompact reloads the swapped files while holding dataFileAccessLock, and for a vacuumed TTL volume that reload re-derived lastModifiedTsSeconds by reading every live needle's append timestamp from the .dat: two random reads per needle, with every read of the volume blocked behind them. The in-memory clock is already current at that point. Every write since the volume loaded moved it, and makeupDiff only replays writes that went through that path. Carry it across the reload instead. This also stops an over-budget scan from falling back to the new .dat's mtime and restarting an expiring volume's TTL at the commit. * volume: carry the append watermark as the TTL clock across a vacuum commit The running append watermark is the clock the reload's recovery scan recomputes, so the commit can keep it directly. Client-supplied needle modified times can run ahead of or behind the append time; keeping lastModifiedTsSeconds itself would let a forged or stale timestamp move expiry through a vacuum, where the scan it replaces used server-side append timestamps. * storage: test that a vacuum commit keeps the append clock A write's client supplied modified time can lie ahead of or behind its append time; the commit must land the TTL clock on the append watermark, the same value the recovery scan would have recomputed. * volume: carry the append watermark as the TTL clock across a vacuum commit Mirrors the Go volume server: the running append watermark is the clock the reload's recovery scan recomputes, so the commit keeps it instead of rescanning live needles under the write lock. * volume: commit carries the last-write append time, not the latest append lastAppendAtNs counts tombstone appends and is reseeded from the .dat tail at every load, so it can sit ahead of the last write -- a delete freshens the commit clock -- or behind it: a restarted vacuumed volume's tail needle is not its newest write, and the commit would move the TTL clock backward into premature expiry. Track lastWriteAppendAtNs instead, bumped only on needle appends and seeded by the recovery scan, so the commit lands the clock on the same live-write maximum the rescan would have recomputed. * volume: rescan at commit when the newest write was deleted lastWriteAppendAtNs can hold a write the index no longer holds, so carrying it extends the TTL clock past what recovery over the compacted index would compute. Remember the key behind the watermark so its tombstone or index rollback can send the reload back through recoverLastModifiedTs, landing on the newest surviving write. * volume: a tombstone retires the write rows beneath it in the last-write scan A needle deleted after the compaction copy leaves its write row followed by a tombstone in the committed index. The reverse scan skipped the tombstone row but then counted the dead write, reseeding the watermark and flag as if it were alive. Track keys whose latest row is a tombstone so their earlier write rows stop counting, and exercise the delete-inside-the-commit-window ordering in the tests. --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
0305e837fd |
iceberg maintenance: keep compacted files prunable (stats, bound order, row groups) (#11654)
* iceberg maintenance: record column statistics on compacted files A compacted file's manifest entry was built with no column_sizes, value_counts, null_value_counts, lower_bounds, upper_bounds or split_offsets, so no reader could skip a compacted file on any predicate. parquet-go already writes exact per-chunk min/max and null counts into the footer; read that footer back after the merge and record it on the data file, bounds as the spec's single-value serialization with string and binary truncated as truncate(16). A column whose bounds cannot be converted exactly gets none, and a statistics failure is logged while the compaction commits anyway: metrics are an optimization, not a correctness requirement. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg maintenance: merge bins in bound order so compacted files stay prunable A bin's files were concatenated in manifest order, or largest-first when a partition was split under the target size, so inputs disjoint on a column came out as outputs that overlapped on it. Order each bin's files by their bounds on one column before merging and split an oversized partition into runs of consecutive files, so every output covers one contiguous range. The column is the first identity field of the table's sort order when it declares one (a descending order sorts by upper bound), otherwise the first schema column every candidate file has bounds for; detection resolves the same order so it plans the bins execution builds. Files without bounds keep the old behavior. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg maintenance: cap compacted files' row groups from table config Neither merge writer set a row-group limit (parquet-go's default is unlimited rows), so every compacted file was a single row group and readers could not skip inside it either. Rows per row group now come from the table's write.parquet.row-group-limit and write.parquet.row-group-size-bytes, defaulting to PyIceberg's 1 048 576 rows and Iceberg's 128 MiB, with the byte size turned into rows from the bin's inputs' compressed bytes per row and a floor of 1 024 rows. Both writers take the cap, and the statistics the entry records list one split offset per row group. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg: keep compaction order eligibility per group, rescue stranded runs The merge order resolved over all candidates, so one oversized or non-Parquet file without bounds disabled ordering for files that could participate. Resolve it per partition group over the eligible entries. Ordered runs too short to merge were dropped entirely. Runs from an inferred bounds order now fall back to size-based packing — ordering is a preference there — while runs under a declared sort order are still left for later passes so the sort contract holds. An explicit write.parquet.row-group-limit is a cap, not a floor: values below the estimate floor are now honored instead of being raised. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg: only use the declared sort order when its bounds are complete An entry without bounds on the sort column sorted to the tail and merged into an output claiming an order it cannot verify. Fall back to bound inference instead of ordering by a later sort field alone. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg maintenance: keep ordered runs when the full repack yields nothing The bestEffort leftover fallback removed the ordered runs before checking whether repacking the whole bin produced any bins, discarding valid compaction work. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
8f80dac30f |
ec: strict_placement option so encode only runs while guarantees hold (#11656)
* ec: strict_placement option so encode only runs while guarantees hold Shard placement during encode was best-effort (PlaceDurabilityFirst): when the cluster could not satisfy the per-disk caps, anti-affinity, replica-placement or per-rack caps, the constraints were relaxed and the volume was encoded anyway, weaker than configured. A strict_placement option on the erasure coding task switches planning to PlaceStrict so the volume's planning fails instead, and the encode is retried when capacity allows the guarantee. Also documents the resilience rule in ec.encode help: a volume survives losing any nodes or racks holding at most parity-shards shards between them, and how -shardReplicaPlacement's rack and node digits bound that loss. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ec: expose strict_placement through the plugin form and persisted task policy The admin UI, the admin.toml maintenance mapping, and the TaskPolicy serialization all dropped the new flag; add the bool field to ErasureCodingTaskConfig, the worker config form, and both conversion directions. * shell: describe shardReplicaPlacement as requested limits, not guarantees ec.encode places shards best-effort, so the configured rack/node caps only bound shard loss when the final placement actually satisfies them. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
30cf53265b |
deps: pin seaweedfs/goexif at v1.0.3 and stop dependabot re-bumping it (#11648)
* deps: pin github.com/seaweedfs/goexif back to v1.0.3 The fork's newest tagged release is v1.0.3; the v2.0.0+incompatible requirement resolved to older, untagged code and breaks isolated builds that fetch it from the proxy. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci: stop dependabot bumping seaweedfs/goexif past its latest tag Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test/kafka: settle goexif at v1.0.3 too The kafka test module recorded v2.0.0+incompatible in its own requires, so MVS kept selecting it over the pinned v1.0.3 in the root module. --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
4d1f49c638 |
filer.sync: sign proxied chunk I/O from the per-side security file (#11645)
* filer.sync: sign proxied chunk I/O from the per-side security file The -a.security / -b.security files were used for gRPC TLS and the HTTPS client but not for jwt.filer_signing, so filer-proxied chunk reads and writes carried a token signed with the process-wide key and failed authorization whenever the two clusters' keys differ. LoadFilerJwtFromFile returns a FilerJwtProvider for each side's file, which FilerSource and FilerSink now accept for proxied chunk reads and writes. With no keys in the file or no flag, both fall back to the process-wide jwt.filer_signing configuration as before. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * replication: use the side filer read key for manifest downloads and fall back per access level Manifest chunk resolution still signed proxied downloads with the process-wide read key, so a source filer requiring its own key 401'd on manifest-bearing files. A side security file that set only one access level also produced empty tokens for the other instead of inheriting the process-wide key, and the side file loader ignored the WEED_ environment overrides the filer itself honors. ResolveChunkManifest/ResolveOneChunkManifest keep their signatures; FilerJwt-aware variants thread the provider down to fetchWholeChunk, which prefers it on proxy URLs. The side loader now applies the same environment precedence and falls back to the process-wide signer per missing access level. * security: verify the configured filer token lifetimes * security: reject negative filer token lifetimes A negative expires_after_seconds reached GenJwtForFilerServer and produced a token with no expiration claim. Also synchronize the Authorization-header capture in the proxy test and restore the prior viper key on cleanup. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
14fdd61aea |
filer: stop the aggregated metadata subscribe loop rescanning an exhausted persisted log (#11644)
* fix(filer): gate the aggregated metadata disk pass on real change A subscriber whose start position is past the end of the local persisted log re-ran the whole persisted-log pass - store listings, file opens, readahead - on every loop iteration. Each iteration is paced only by the shortest wake (the 20ms hold floor on a busy watermark), so one parked subscriber kept a full CPU core busy for the life of the stream. The aggregated loop now mirrors the local loop's gate: the disk pass runs on the first pass and afterwards only when something it cannot miss changed - a local flush landed, the peers' flush low-watermark advanced (more content admitted, or new files in a shared store), the cursor moved, or a disk hold is pending (the ring read that follows an empty pass parks internally, so skipping there would strand a held entry). Regression test: a subscriber parked past the persisted-log tail holds the listing rate near zero and still delivers once peers report progress. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: re-arm the aggregated disk pass on unobserved change Review found three staleness classes the gate could not see: the flush low-watermark only catching rises (a joining peer lowers the minimum and invalidates an earlier pass's proof), a peer past the minimum landing a file without moving it, and a chunk subscriber's refs-stop bound advancing with wall time. Re-read when the low-watermark moves in either direction, when the chunk listing bound admits more files, and on a slow re-probe cadence for files no watermark can signal. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: unwind the parked ring read so the disk re-probe runs, and re-read on cursor rewinds A caught-up subscriber parks inside LoopProcessLogData's wait loop, so the re-probe interval in the outer disk gate could never elapse there; the callback now unwinds the read once the cadence is due so the gate re-evaluates. The cursor trigger also needs to notice rewinds, not just advances, since ResumeFromDiskError moves the cursor backward. --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
49c25890ad | docs: regenerate star history chart | ||
|
|
56fbc3deb5 |
s3api: align S3/IAM error responses with AWS (#11632)
* s3err: add InvalidArgument and AuthorizationHeaderMalformed codes Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: answer unrecognized bucket PUT sub-resources with 501 A PUT on a bucket carrying an unrecognized query (logging, metrics, intelligent-tiering, ...) fell through to the bare CreateBucket route and returned BucketAlreadyOwnedByYou or re-created the bucket. AWS answers these with NotImplemented. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: reject malformed copy-source and multipart PUT parameters A malformed X-Amz-Copy-Source or a non-numeric partNumber fell through to the plain PutObject route and stored the body as a regular object. Answer them with InvalidArgument-class errors instead of writing data. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: verify x-amz-content-sha256 against the streamed body A PUT carrying a hex or base64 payload hash now streams through a verifier that reports a mismatch once the stream is exhausted, instead of storing an object that does not match its declared hash. The error is deferred so intermediate reads that drop (n>0, err) results cannot silently swallow it. Sentinel values (unsigned/streaming payloads) remain exempt and malformed values fail fast. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: map truncated SigV4 headers to the error for the missing field AWS answers an Authorization header missing Credential= with InvalidArgument and one missing or malformed Signature= with AuthorizationHeaderMalformed, instead of a generic MissingFields. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: answer IAM/STS failures in the query-protocol envelope Embedded IAM and STS routes now report authentication, form-parse and authorization failures with the IAM ErrorResponse body instead of the S3 Error envelope, so IAM SDK clients can parse them. Requests signed for s3 keep the S3 envelope, keyed off the credential scope. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: return 403 AccessDenied when the request has no Date Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: answer throttling rejections as SlowDown ErrTooManyRequest and ErrRequestBytesExceed reported made-up codes; AWS serves these throttling rejections as SlowDown. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: reject versionId requests on buckets that never had versioning GET, HEAD and DELETE carrying a non-empty versionId on an unversioned bucket now fail with InvalidArgument instead of being answered as a plain object request. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: answer DeleteObjects over 1000 keys with MalformedXML Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: InvalidArgument for non-numeric or out-of-range part numbers partNumber=abc, 0 and >10000 all resolve to InvalidArgument, matching AWS, instead of InvalidPart. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iam: correct LimitExceeded, InvalidAction and ServiceFailure mappings LimitExceeded is a conflict (409), an unknown Action is InvalidAction (404) rather than NotImplemented, and internal failures report the IAM receiver fault type. Applies to both the embedded IAM endpoint and the standalone iamapi server. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iam: refuse DeleteUser while access keys remain Deleting a user with live credentials orphaned its access keys; AWS answers DeleteConflict until they are removed first. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test/s3: add S3/IAM error-response compatibility harness * s3: return InvalidArgument for malformed x-amz-content-sha256 A header value that decodes to neither 32-byte hex nor base64 is a malformed argument, not a hash mismatch. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: stop spawned mini when readiness times out A slow-starting server otherwise survives the failure path and keeps the S3 port occupied for the next run. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: register atexit cleanup before setup A failed setup previously skipped cleanup, leaking the bucket and IAM user on persistent servers. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3api: reject unrouted subresources on the DELETE bucket catch-all PutBucketHandler gained the same guard when the route-level check moved into the handlers; DeleteBucketHandler was missed, so an authorized DELETE /bucket?logging could delete the bucket instead of answering NotImplemented. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: tolerate unset fixture variables in cleanup and cover DELETE ?logging Cleanup now runs its IAM/multipart steps only when setup reached them, so an early setup failure still removes the bucket. Added a DELETE bucket-subresource case asserting NotImplemented and bucket survival. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: fail a case when its side-effect check reports a regression A non-empty check note now fails the case, so a deleted bucket or an object created by a malformed request cannot slip through behind a passing status check. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
079a7d8ba3 |
s3: give each prefix its own hidden-prefix probe budget (#11626)
* s3: give each prefix its own hidden-prefix probe budget The shared cursor.probedEntries counter in dirHoldsOnlyHiddenEntries was exhausted by one large all-deleted subtree, causing every later prefix in the same request to be treated as visible. The fix allocates a fresh budget (hiddenProbePerPrefixBudget, default 1000) for each top-level call and threads it down to recursive calls via a pointer, so cross-prefix budget bleed is impossible. Fixes #10847. * s3: keep 10000 probe budget, now per prefix Restore hiddenProbeBudget = 10000 as a package const (not a mutable var) applied per-prefix instead of per-request. The override hook moves to an unexported probeBudget field on ListingCursor; zero means use the package default. Tests set probeBudget: 3 on the cursor, keeping them fast without touching package state. Also revert the unrelated uint32 cast on the ListEntries Limit field. * s3: cap total hidden-prefix probe work per listing request The per-prefix budget resets for every candidate prefix, but deleted prefixes do not spend maxKeys, so a page can walk an unbounded number of them. Keep a request-wide probedEntries ceiling (hiddenProbeTotalBudget, 10x the per-prefix budget) so total probe work stays bounded. * s3api: pin the probe-budget cutoff and the request-wide ceiling --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
0bcebf708c |
ecbalancer: cap total shards per rack in Plan (#11623)
* ecbalancer: cap total shards per rack in Plan Plan caps data and parity per rack separately (ceil(data/racks) and ceil(parity/racks)), so with 10+4 over 8 racks a rack can legally hold 2 data + 1 parity. When a rack is one disk, losing two such racks loses 6 of 14 shards and the volume can't be read. The cross-rack phase now also caps each rack's TOTAL shards of a volume, sized with Place's rackTotalCap: ceil(shards/racks) unless the racks lack room, counting a rack's own shards of the volume as room since Plan can move them. - A rack above the cap sheds parity until it fits. Those shards may go to a data-bearing rack, and when no rack is under the parity cap they fall back to a rack under the total cap. #11438's non-overflow candidates keep moving only to data-free racks. - No cross-rack move lands on a rack at the cap. - A rack above the cap triggers balancing regardless of the imbalance threshold. - The fallback applies only while the source rack is above the cap; otherwise the next Plan moves the shard back. - Options.RackTotalCapRaised reports volumes whose cap had to be raised above the even share; the worker and the shell log it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ecbalancer: finish the rack cap in one Plan, count SameRackCount room Review follow-ups. - The data pass can run out of destinations before the parity pass frees slots elsewhere, which left a data-heavy rack above the cap after one Plan (a one-shot shell balance stops there). The cross-rack phase now repeats while a rack is above the cap and the last round moved something. A shard moves at most once per plan, since each move runs as its own task. - planRackTotalCap bounds each node's room by what SameRackCount still allows, as Place does. Counting the raw free slots sized the cap too low, so RackTotalCapRaised missed volumes it should report. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
576837f1b2 |
fix(s3): make self-heal pointer persist CAS-bound against concurrent writers (#11627)
* fix(s3): make self-heal pointer persist CAS-bound against concurrent writers Follow-up to #11618: pointerless reads of a slash key whose regular-path entry is a physical parent (or a bare-key object) now fall through to healStaleLatestVersionPointer, which rescans .versions and persists a repaired pointer. The persist was an unconditional upsert off the pre-scan snapshot, so a PUT or delete that atomically advanced the pointer on the owner filer while the heal was rescanning could be rolled back, making older content or ACLs current again. Mirror the CAS discipline clearStaleLatestVersionPointer already applies: re-fetch the live .versions entry, require its pointer fields to still match the ones the heal observed, and abandon the persist (still returning the rescanned entry) when a concurrent writer has moved them. Write the live Extended map so concurrently updated fields are preserved. * fix(s3): close the check-then-act window in the self-heal pointer persist The CAS re-fetch added in the previous commit narrows the race but leaves a gateway-side window: after the live .versions entry is re-read and the pointer compared, the repair is still written back through an unconditional RPC, so a PUT or delete committing between the re-fetch and the persist still ends up rolled back by the stale repair. Bind the persist to the live image the heal just re-read with an IF_ENTRY_EQUAL precondition, the same discipline routedSelfCopy applies to stale self-copies: the filer evaluates the condition under the entry's path lock and conditional writes route to the owner filer, so a writer committing inside the window fails the precondition and the winner's pointer stands. FailedPrecondition and NotFound are authoritative replies and are not replayed by the failover layer. The test now also covers a writer committing during the persist, which reverts the pointer on the previous unconditional write-back. * s3api: CAS-bind the stale-pointer clear against concurrent writers The pointer clear re-read the live .versions entry and then wrote it back unconditionally through mkFile, so a writer committing between the re-fetch and the persist was rolled back to a cleared pointer. Persist through the same IF_ENTRY_EQUAL conditional update as the repair path. * s3api: test the CAS contract on the stale-pointer clear * s3api: never clear a pointer the clear did not observe as stale The CAS clear skipped its live-pointer match when the caller's snapshot carried an empty latest-version id, so a writer promoting a version between the clear's rescan and its re-fetch had the fresh pointer CAS-cleared away (expected = the writer's own live entry), briefly making the just-written version appear absent. With an empty observed id, reaching the persist at all implies a concurrent promotion (an idle key short-circuits as already-clear), so make the pointer match unconditional and abort instead. Extend TestClearStaleLatestVersionPointerConcurrentWriter with pointerless-snapshot cases: a post-rescan promotion must survive, and an idle pointerless key must short-circuit as already-clear. The fake filer's proto round-trip drops empty Extended maps, so the snapshot is padded the way real callers do. --------- Co-authored-by: zhaoyuchen <yc.zhao@yinzon.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
3288d90b21 | docs: regenerate star history chart | ||
|
|
199d78539a |
fix(cosi): add get/list/watch on bucketclasses to enable static provi… (#11620)
* fix(cosi): add get/list/watch on bucketclasses to enable static provisioning * fix(cosi): bind provisioner deployment to the chart-created service account The deployment referenced a release-prefixed service account name while the chart creates and binds "seaweedfs-objectstorage-provisioner", so the RBAC grants never reached the provisioner pod. --------- Co-authored-by: Jonas Onshuus <jons@dips.no> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
ec261c5fbc |
S3: quote ETag in CopyObject and UploadPartCopy XML responses (#11624)
* s3api: extract quoteETag from setEtag Consolidate ETag quoting so the XML response builders can share it. * s3api: quote ETag in CopyObjectResult XML AWS returns the ETag quoted in the copy result body, matching the ETag header. Fixes seaweedfs/seaweedfs#11622 * s3api: quote ETag in CopyPartResult XML UploadPartCopy returned the raw ETag in the XML body while the response header and other APIs return it quoted. Fixes seaweedfs/seaweedfs#11622 * s3api: test ETag quoting in copy responses * s3api: assert quoted ETag on the wire in copy response tests * s3api: use strconv.Quote in quoteETag Addresses CodeQL 'potentially unsafe quoting' on string concatenation. |
||
|
|
cac6cd4b16 |
build(deps): bump rustls from 0.23.43 to 0.23.45 in /seaweed-worker (#11617)
Bumps [rustls](https://github.com/rustls/rustls) from 0.23.43 to 0.23.45. - [Release notes](https://github.com/rustls/rustls/releases) - [Changelog](https://github.com/rustls/rustls/blob/main/CHANGELOG.md) - [Commits](https://github.com/rustls/rustls/compare/v/0.23.43...v/0.23.45) --- updated-dependencies: - dependency-name: rustls dependency-version: 0.23.45 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
d829275de6 |
fix(s3): recheck cache metadata and validate null objects (#11618)
* fix(s3): recheck cache metadata and read null directory markers Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * fix(s3): preserve null object identity and bound test cache Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * s3: trim comments on the anonymous read cache path Signed-off-by: Chris Lu <chris.lu@gmail.com> --------- Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> Signed-off-by: Chris Lu <chris.lu@gmail.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
23893eb378 |
volume: return error instead of panicking when .dat open fails (#11619)
* volume: return error instead of panicking when .dat open fails When backend.OpenVolumeFile returns an error (e.g. disk below -minFreeSpace), dataFile is nil. Calling backend.NewDiskFile(nil) immediately after caused a nil-pointer panic in f.Stat()/f.Name(). Move the existing error check to run right after OpenVolumeFile, before NewDiskFile is called, so the error is returned cleanly. Fixes seaweedfs/seaweedfs#11615 * volume: share .dat load error handling Extract datFileLoadError helper so open and create paths share one check. * volume: trim dat-open-fail test comments * volume(rust): cover unopenable .dat load path --------- Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
32e77ff980 |
Secure Weed Mini Admin Listeners by Default (#11613)
* securing admin Signed-off-by: Subhadeep Maity <smaity@slb.com> * updated readme Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> * fixed pr comments Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> * docs: tidy weed mini admin bind notes Drop the new single-entry CHANGELOG.md since changes are documented via GitHub releases, and rewrap the README paragraph to match the surrounding one-line style without self-referential issue/PR links. * review comments Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> --------- Signed-off-by: Subhadeep Maity <smaity@slb.com> Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> Co-authored-by: Subhadeep Maity <smaity@slb.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
de75450655 |
build(deps): bump github.com/apache/iceberg-go from 0.6.1-0.20260817192109-c2105090c9e2 to 0.7.0 (#11609)
* build(deps): bump github.com/apache/iceberg-go Bumps [github.com/apache/iceberg-go](https://github.com/apache/iceberg-go) from 0.6.1-0.20260817192109-c2105090c9e2 to 0.7.0. - [Release notes](https://github.com/apache/iceberg-go/releases) - [Commits](https://github.com/apache/iceberg-go/commits/v0.7.0) --- updated-dependencies: - dependency-name: github.com/apache/iceberg-go dependency-version: 0.7.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * iceberg: write delete manifests with NewManifestWriter iceberg-go 0.7.0 validates that delete entries only land in delete-content manifests, so the WriteManifest + byte-level content patch workaround no longer works: WriteManifest always creates a data-content writer and rejects delete entries outright. Use NewManifestWriter with WithManifestWriterContent instead and dispatch entries through Add/Existing/Delete by status. The returned ManifestFile now carries writer-computed counts, partitions and min-sequence-number, so the manual ManifestFile rebuild is dropped. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
b14dd1cee4 |
helm: drop fromToml dependency in security-configmap.yaml (fixes #11611) (#11614)
* helm: drop fromToml dependency in security-configmap.yaml (fixes #11611) fromToml was added to Helm in v3.17.0 (helm/helm#12026, merged 2024-09-12, one day after v3.16.0 was cut). This chart declares no minimum Helm version (no Chart.yaml kubeVersion, nothing in the README), and the call in security-configmap.yaml:21 is an unconditional *parse*-time failure on Helm < v3.17.0 - Go's text/template parses a file's entire body before evaluating any {{if}}, so this breaks the chart (any topology, any values) even when securityConfigEnabled is false and the ConfigMap would render nothing. Replaces the fromToml-based dig lookup with a small regex-based helper (seaweedfs.existingTomlKey) that reads the same four "key = ..." JWT signing-key values out of a previously-rendered security.toml, preserving the existing fallback-to-random behavior exactly. Verified: - helm lint (v3.16.3 and v4.3.0): clean - helm template with chart defaults: byte-identical output to the unpatched chart rendered via Helm v4 (which has fromToml) - the disabled/default path is untouched - helm template with security enabled, no prior ConfigMap: identical structure to the unpatched chart (helm v4), modulo the expected random key - Real helm install + helm upgrade round trip (live lookup, since "helm template" never evaluates lookup, even under the original fromToml code): the JWT signing key is identical across both releases - confirms key persistence across upgrades is preserved, not just "renders without erroring" - helm template with chart defaults, Helm v3.16.3: previously failed with a parse error naming fromToml as undefined; now renders successfully Fixes #11611. * helm: harden existingTomlKey against commented key lines and CRLF Addresses two review findings from greptile-apps on PR #11614: - The key-line regex matched the first "key = ..." anywhere in the section block, including a commented-out "# key = ..." line, which would shadow a real active key on a hand-edited or otherwise non-chart-generated ConfigMap. Anchored to line start with the Go regexp multiline flag ((?m)^key...), which a line starting with "#" cannot match. - The section-header match required an exact "]\n", so a ConfigMap with CRLF line endings would fail to match the block at all and regenerate the key instead of reusing it. Changed to "]\r?\n". Also adds a CI test ("Verify JWT signing key persistence across upgrades") exercising all of this end to end with a real helm install -> edit the live ConfigMap -> helm upgrade cycle, matching the existing "Verify SFTP host key secret lifecycle" test's shape: both edge cases are reproduced against a real ConfigMap and asserted on the post-upgrade rendered security.toml. Verified locally (same commands as the new CI step) against a real cluster before pushing. * helm: preserve JWT keys across supported TOML layouts * ci: use setup-python interpreter for JWT upgrade checks * helm: preserve keys under quoted TOML section headers * helm: ignore unrelated quoted TOML section headers |
||
|
|
3a65365f6e |
build(deps): bump org.apache.spark:spark-core_2.12 from 3.5.7 to 3.5.8 in /test/java/spark (#11616)
build(deps): bump org.apache.spark:spark-core_2.12 in /test/java/spark Bumps org.apache.spark:spark-core_2.12 from 3.5.7 to 3.5.8. --- updated-dependencies: - dependency-name: org.apache.spark:spark-core_2.12 dependency-version: 3.5.8 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
cb053e601a |
build(deps): bump github.com/tidwall/gjson from 1.18.0 to 1.19.0 (#11607)
Bumps [github.com/tidwall/gjson](https://github.com/tidwall/gjson) from 1.18.0 to 1.19.0. - [Commits](https://github.com/tidwall/gjson/compare/v1.18.0...v1.19.0) --- updated-dependencies: - dependency-name: github.com/tidwall/gjson dependency-version: 1.19.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
8ebef03509 |
helm: expose s3 update strategy, lifecycle, and termination grace period (#11602)
The standalone s3 Deployment hardcoded terminationGracePeriodSeconds and offered no way to set a container lifecycle or the Deployment strategy, so operators could not add a preStop delay to drain endpoints before SIGTERM or hold maxUnavailable at 0 during rollouts. Add s3.updateStrategy, s3.lifecycle and s3.terminationGracePeriodSeconds. Defaults render the same manifest as before. |
||
|
|
c739e5cf78 |
fix(s3): drain in-flight requests on SIGTERM in standalone weed s3 (#11603)
Standalone weed s3 only shut its servers down when shutdownCtx was set, which only weed mini does. On SIGTERM the interrupt hooks ran and the process exited with requests still in flight, so clients saw connection resets during rolling restarts. Register an interrupt hook that drains every S3 HTTP(S) listener (including the local and unix socket ones, and the Iceberg and Lance servers) for up to 15s alongside a bounded gRPC GracefulStop, then closes the S3 API server, reusing the filer's shutdown helper. Serve exits now join that shutdown so the process does not exit mid-drain, and the secondary listeners tolerate http.ErrServerClosed instead of exiting fatally. Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
c0a751072d |
volume server: apply Range to chunk manifests and forward raw headers when proxying (#11538)
* volume server: read GET/HEAD needles off the store lock, and only once The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stream needle chunks without the store lock StreamingBody::poll_frame took store.read() and find_volume for every chunk to compare the volume's compaction revision, dup'd the source handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the stream also holds a data-file read lease for its whole life, while a writer waits for that lease under store.write(): the next chunk's store.read() then waits for the writer and the writer for the stream. The per-chunk re-lookup was also wrong. The stream reads a handle opened at plan time, which pins the .dat inode the offset was resolved against; a vacuum commit renames a new file over .dat and leaves that inode untouched. The re-looked-up offset belongs to the new file but was read from the old inode, so a stream whose needle a vacuum moved ended in a checksum error. The pinned offset stays valid, so the check, and with it every store access, is dropped, along with the now unused re_lookup_needle_data_offset and the revision fields of the read plan. The source is shared as an Arc instead of dup'd per chunk, and the chunk buffer is a BytesMut that the blocking read hands back with its result, so its allocation is reclaimed once the previous frame has been written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: split get_or_head_handler_inner into phases get_or_head_handler_inner was a ~650-line function. Its middle resolved the needle and set five mutable flags (stream_info, can_stream, can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that three if-let reply paths then re-tested, each re-checking stream_info. It is now a 126-line orchestrator over named phases: reject_read_jwt, proxy_missing_volume, wait_for_download_slot, parse_read_request, read_ec_needle / read_volume_needle, etag_and_last_modified, not_modified_response, read_response_headers, and the reply phases stream_response, head_from_meta_response, range_from_source_response, buffered_payload and buffered_response. The read phases return a ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource, Buffered) carries the NeedleStreamInfo only on the variants that use it, so the reply is one match instead of three flag checks. Pure refactor: every status code, header and header order, error text, metric increment, lock and data-file lease scope, spawn_blocking boundary and side-effect order is unchanged. Phases that can end the request return ControlFlow<Response, T>. A Range header that is not visible ASCII still falls through to the buffered path, as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a needle stream once its volume becomes unavailable Taking the store lock out of StreamingBody also dropped its per-chunk unavailable_error() check. With -hasSlowRead a writer can take the data-file lease between chunks, fail its fsync and its truncate, and mark the volume unavailable; the stream then kept serving the rest of the needle from its pinned handle. The volume's io_unavailable reason is now an Arc-shared leaf mutex that the read plan hands to the stream. Each chunk checks it under its data-file lease, where the writer marks it, and fails with the same "volume is unavailable: <reason>" error the old check returned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: read a non-ASCII or empty Range header as Go does A Range value with a byte >= 0x80 (obs-text, which hyper accepts) failed HeaderValue::to_str at both range gates. For a needle served as stored the handler had already chosen a meta-only read, so it fell through to the buffered path with no payload and answered 200 with an empty body; a compressed or EC needle answered 200 with the full body. Go's parseRange fails on a byte it can neither trim nor parse and answers 416 "invalid range", and trims Unicode whitespace such as NBSP into a normal 206. An empty Range value was also a 200 with an empty body, where Go sends the whole payload. Read Range once with from_utf8_lossy, dropping an empty value, and hand that one value to the read plan and to both range gates. A replaced byte never parses, so it is a 416; str::trim trims the same Unicode whitespace as strings.TrimSpace. A range read from the data file now always answers itself instead of falling through with an empty needle. The buffered path answers HEAD before it looks at Range, as Go's writeResponseContent does, so an EC HEAD with a Range is a 200 with the full length. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: apply Range to chunk manifests and forward raw headers when proxying A GET of a chunk manifest assembled the object and always answered 200 with the whole body, ignoring Range. Go serves the expanded manifest through writeResponseContent, which answers HEAD first and then hands Range to ProcessRangeRequest: 206 for one range, multipart/byteranges for several, 416 for an unsatisfiable or unparsable one. try_expand_chunk_manifest now returns the assembled body and headers, and the caller answers through buffered_response, the same helper the buffered needle path uses. A proxied read forwarded a request header only if HeaderValue::to_str succeeded, so a Range with an obs-text byte was dropped and the target answered 200 with the full body. Go copies every header value as is. Forward the raw HeaderValue for every header. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: reset Content-Type on range errors as net/http does Go's http.Error sets text/plain and nosniff unconditionally; keeping the needle's MIME type on a 416 mislabels the error body. Match it. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: count manifest downloads toward the limit Expanded manifests hard-coded track_download off, so chunked-file reads bypassed in-flight byte accounting while every other path honored the download limiter. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
52c6df3bec |
fix(filer): leave the lock ring before stopping gRPC on shutdown (#11604)
A draining filer stayed in the master's lock ring until its process exited, while gRPC GracefulStop was already refusing new connections. S3 gateways and peer filers kept routing object-write locks and owner-routed writes to it for the whole graceful-stop window. On shutdown the filer now sends leave_lock_ring on its open KeepConnected stream. The master removes it from the lock ring only; it stays a cluster member so peers keep following its metadata log through the drain. Once the ring update without it arrives, the filer has already transferred its locks to the new owners, and it keeps serving through the prior-owner window (plus a second for peers that apply the update later) before gRPC and HTTP begin draining. The whole leave is bounded at 10s so slow lock transfers or a stuck stream cannot hold up the drain. A lone filer, or a master that ignores the message, falls back to the previous behavior. |
||
|
|
aa5b337716 |
fix(s3): honor object ACLs for anonymous GET and HEAD (#11605)
* fix(s3): honor object ACLs for anonymous reads Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * s3: drop unsigned session tokens on deferred anonymous reads An unsigned GET/HEAD carrying X-Amz-Security-Token is anonymous, not a session request; strip the token before authentication and authorization so it cannot influence identity or policy evaluation. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
abdb95dc87 |
build(deps): bump golang.org/x/time from 0.15.0 to 0.16.0 (#11608)
Bumps [golang.org/x/time](https://github.com/golang/time) from 0.15.0 to 0.16.0. - [Commits](https://github.com/golang/time/compare/v0.15.0...v0.16.0) --- updated-dependencies: - dependency-name: golang.org/x/time dependency-version: 0.16.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
076fd24186 |
filer: bound metadata log flush retries during shutdown (#11596)
* filer: bound metadata log flush retries during shutdown On SIGTERM the filer could hang forever in Shutdown: the final LocalMetaLogBuffer flush retries appendToFile indefinitely, and with the master already down each AssignVolume attempt just kept failing. WaitForShutdown never returned, the interrupt hook never reached os.Exit, and the half-dead filer kept its ports bound. Thread a context through appendToFile/assignAndUpload and switch to a 15s-bounded context once the filer is stopping: the flush abandons with a log line instead of retrying forever. Normal operation keeps the unbounded retry so no metadata is dropped while running. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: share one shutdown deadline across all pending meta log flushes Review feedback on the per-flush timeout: a deadline armed at flush start could already be expired when shutdown arrived, the first append of a flush still ran unbounded, each queued window got a fresh budget (16 windows * 15s), and an abandoned window still advanced the flushed watermark as if it had landed. Rework to a single shared flush context on the Filer, cancelled once by Shutdown via AfterFunc. Every append - the in-flight one and every queued window - observes the same deadline, so the whole drain is bounded at 15s. A flush that gives up reports its dropped bytes through the new LogBuffer.NoteFlushDropped, and loopFlush then skips the offset/timestamp advance and subscriber notifications for that window. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: bound append attempts and commit uploaded pieces detached Store operations check then drop request cancellation, so a stalled backend could still hold flushFn past the shutdown deadline; run each append attempt on its own goroutine and give up on it at the deadline. Once a piece is uploaded, commit its entry on a detached context so the expired deadline cannot strand the chunk. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: keep shutdown flushes synchronous Detaching the append attempt let flushFn return while the goroutine still held the pooled flush buffer and could commit after the metadata store closed; abandonment is only safe for the cancelable assign/upload phase, which the shared flush context already bounds. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
90eb4ec091 |
build(deps): bump github.com/hashicorp/raft-boltdb/v2 from 2.3.1 to 2.4.2 (#11606)
build(deps): bump github.com/hashicorp/raft-boltdb/v2 Bumps [github.com/hashicorp/raft-boltdb/v2](https://github.com/hashicorp/raft-boltdb) from 2.3.1 to 2.4.2. - [Release notes](https://github.com/hashicorp/raft-boltdb/releases) - [Commits](https://github.com/hashicorp/raft-boltdb/compare/v2.3.1...v2.4.2) --- updated-dependencies: - dependency-name: github.com/hashicorp/raft-boltdb/v2 dependency-version: 2.4.2 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
e710cfc4b0 |
build(deps): bump github.com/rabbitmq/amqp091-go from 1.14.0 to 1.15.0 (#11610)
Bumps [github.com/rabbitmq/amqp091-go](https://github.com/rabbitmq/amqp091-go) from 1.14.0 to 1.15.0. - [Release notes](https://github.com/rabbitmq/amqp091-go/releases) - [Changelog](https://github.com/rabbitmq/amqp091-go/blob/main/CHANGELOG.md) - [Commits](https://github.com/rabbitmq/amqp091-go/compare/v1.14.0...v1.15.0) --- updated-dependencies: - dependency-name: github.com/rabbitmq/amqp091-go dependency-version: 1.15.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
cf144d5eb2 | docs: regenerate star history chart | ||
|
|
825c3dff8b |
volume server: fetch only the chunks a manifest Range needs, like Go (#11546)
* volume server: read GET/HEAD needles off the store lock, and only once The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stream needle chunks without the store lock StreamingBody::poll_frame took store.read() and find_volume for every chunk to compare the volume's compaction revision, dup'd the source handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the stream also holds a data-file read lease for its whole life, while a writer waits for that lease under store.write(): the next chunk's store.read() then waits for the writer and the writer for the stream. The per-chunk re-lookup was also wrong. The stream reads a handle opened at plan time, which pins the .dat inode the offset was resolved against; a vacuum commit renames a new file over .dat and leaves that inode untouched. The re-looked-up offset belongs to the new file but was read from the old inode, so a stream whose needle a vacuum moved ended in a checksum error. The pinned offset stays valid, so the check, and with it every store access, is dropped, along with the now unused re_lookup_needle_data_offset and the revision fields of the read plan. The source is shared as an Arc instead of dup'd per chunk, and the chunk buffer is a BytesMut that the blocking read hands back with its result, so its allocation is reclaimed once the previous frame has been written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: split get_or_head_handler_inner into phases get_or_head_handler_inner was a ~650-line function. Its middle resolved the needle and set five mutable flags (stream_info, can_stream, can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that three if-let reply paths then re-tested, each re-checking stream_info. It is now a 126-line orchestrator over named phases: reject_read_jwt, proxy_missing_volume, wait_for_download_slot, parse_read_request, read_ec_needle / read_volume_needle, etag_and_last_modified, not_modified_response, read_response_headers, and the reply phases stream_response, head_from_meta_response, range_from_source_response, buffered_payload and buffered_response. The read phases return a ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource, Buffered) carries the NeedleStreamInfo only on the variants that use it, so the reply is one match instead of three flag checks. Pure refactor: every status code, header and header order, error text, metric increment, lock and data-file lease scope, spawn_blocking boundary and side-effect order is unchanged. Phases that can end the request return ControlFlow<Response, T>. A Range header that is not visible ASCII still falls through to the buffered path, as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a needle stream once its volume becomes unavailable Taking the store lock out of StreamingBody also dropped its per-chunk unavailable_error() check. With -hasSlowRead a writer can take the data-file lease between chunks, fail its fsync and its truncate, and mark the volume unavailable; the stream then kept serving the rest of the needle from its pinned handle. The volume's io_unavailable reason is now an Arc-shared leaf mutex that the read plan hands to the stream. Each chunk checks it under its data-file lease, where the writer marks it, and fails with the same "volume is unavailable: <reason>" error the old check returned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: read a non-ASCII or empty Range header as Go does A Range value with a byte >= 0x80 (obs-text, which hyper accepts) failed HeaderValue::to_str at both range gates. For a needle served as stored the handler had already chosen a meta-only read, so it fell through to the buffered path with no payload and answered 200 with an empty body; a compressed or EC needle answered 200 with the full body. Go's parseRange fails on a byte it can neither trim nor parse and answers 416 "invalid range", and trims Unicode whitespace such as NBSP into a normal 206. An empty Range value was also a 200 with an empty body, where Go sends the whole payload. Read Range once with from_utf8_lossy, dropping an empty value, and hand that one value to the read plan and to both range gates. A replaced byte never parses, so it is a 416; str::trim trims the same Unicode whitespace as strings.TrimSpace. A range read from the data file now always answers itself instead of falling through with an empty needle. The buffered path answers HEAD before it looks at Range, as Go's writeResponseContent does, so an EC HEAD with a Range is a 200 with the full length. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: apply Range to chunk manifests and forward raw headers when proxying A GET of a chunk manifest assembled the object and always answered 200 with the whole body, ignoring Range. Go serves the expanded manifest through writeResponseContent, which answers HEAD first and then hands Range to ProcessRangeRequest: 206 for one range, multipart/byteranges for several, 416 for an unsatisfiable or unparsable one. try_expand_chunk_manifest now returns the assembled body and headers, and the caller answers through buffered_response, the same helper the buffered needle path uses. A proxied read forwarded a request header only if HeaderValue::to_str succeeded, so a Range with an obs-text byte was dropped and the target answered 200 with the full body. Go copies every header value as is. Forward the raw HeaderValue for every header. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: fetch only the chunks a manifest Range needs, like Go A ranged GET of a chunk manifest fetched every chunk, assembled the whole object and then sliced it, so reading a few bytes of a large object cost a read of all of it, and a 416 still fetched everything. Go serves a manifest through ChunkedFileReader, which seeks to each range and reads only the chunks under it. For a GET with a Range, try_expand_chunk_manifest now parses the ranges against the manifest size, fetches only the chunks whose declared window overlaps one of them (none when the reply carries no body), and answers through handle_range_request_with, the reader-based core that handle_range_request now wraps, so 206/416/multipart stay one code path. The reader replays assembly: chunks clamped as before, later chunks over earlier ones, zeros in gaps. HEAD, no-Range GETs and GETs that crop or resize an image still assemble the whole object. A missing chunk outside the requested ranges no longer turns a ranged GET into a 500, as in Go; a missing chunk inside them still does, before any headers are sent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: reword a comment codespell flags Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: keep only range-covered bytes of fetched manifest chunks A ranged GET retained every overlapping chunk's full contents; 1,000 overlapping 8 MiB chunks could pin ~8 GiB for a one-byte response. Clip each fetched chunk to the bytes the requested ranges can actually read, preserving the later-chunks-overwrite and zero-fill-gap semantics. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: bucket ranged manifest parts by range Serving a multipart range scanned every retained part. Bucket the kept intersections by their range so one range only reads its own parts. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
a7a590474d |
kafka: honor notification.kafka.event_types (#11601)
* kafka: honor notification.kafka.event_types Kafka published every filer event and ignored the filter the webhook notifier already uses. Co-authored-by: Cursor <cursoragent@cursor.com> * notification: share event-type classification between queues Kafka duplicated the webhook's event classification verbatim; move it to the notification package so the two queues cannot drift. Webhook keeps its typed eventType wrappers over the shared helpers. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
8c67b75190 |
image: add an optional public image processing gateway (#11593)
* image: add an optional public image processing gateway * image: fix representation metadata and processing bounds * image: restrict passthrough to non-executable media types Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * image: tighten source media-type validation Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * image: write passthrough body on the inner response writer CodeQL still flagged the passthrough write: the content type was set on the wrapper while the body reached w.ResponseWriter, so the validated header could not be associated with the write. Set headers and copy the body on the same inner writer. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * image: serve processed output on the inner response writer * image: reject XML source types and unsafe conditional metadata --------- Co-authored-by: zhaoyuchen <yc.zhao@yinzon.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
7a19961074 |
volume server: read a non-ASCII or empty Range header as Go does (#11534)
* volume server: read GET/HEAD needles off the store lock, and only once The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stream needle chunks without the store lock StreamingBody::poll_frame took store.read() and find_volume for every chunk to compare the volume's compaction revision, dup'd the source handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the stream also holds a data-file read lease for its whole life, while a writer waits for that lease under store.write(): the next chunk's store.read() then waits for the writer and the writer for the stream. The per-chunk re-lookup was also wrong. The stream reads a handle opened at plan time, which pins the .dat inode the offset was resolved against; a vacuum commit renames a new file over .dat and leaves that inode untouched. The re-looked-up offset belongs to the new file but was read from the old inode, so a stream whose needle a vacuum moved ended in a checksum error. The pinned offset stays valid, so the check, and with it every store access, is dropped, along with the now unused re_lookup_needle_data_offset and the revision fields of the read plan. The source is shared as an Arc instead of dup'd per chunk, and the chunk buffer is a BytesMut that the blocking read hands back with its result, so its allocation is reclaimed once the previous frame has been written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: split get_or_head_handler_inner into phases get_or_head_handler_inner was a ~650-line function. Its middle resolved the needle and set five mutable flags (stream_info, can_stream, can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that three if-let reply paths then re-tested, each re-checking stream_info. It is now a 126-line orchestrator over named phases: reject_read_jwt, proxy_missing_volume, wait_for_download_slot, parse_read_request, read_ec_needle / read_volume_needle, etag_and_last_modified, not_modified_response, read_response_headers, and the reply phases stream_response, head_from_meta_response, range_from_source_response, buffered_payload and buffered_response. The read phases return a ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource, Buffered) carries the NeedleStreamInfo only on the variants that use it, so the reply is one match instead of three flag checks. Pure refactor: every status code, header and header order, error text, metric increment, lock and data-file lease scope, spawn_blocking boundary and side-effect order is unchanged. Phases that can end the request return ControlFlow<Response, T>. A Range header that is not visible ASCII still falls through to the buffered path, as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a needle stream once its volume becomes unavailable Taking the store lock out of StreamingBody also dropped its per-chunk unavailable_error() check. With -hasSlowRead a writer can take the data-file lease between chunks, fail its fsync and its truncate, and mark the volume unavailable; the stream then kept serving the rest of the needle from its pinned handle. The volume's io_unavailable reason is now an Arc-shared leaf mutex that the read plan hands to the stream. Each chunk checks it under its data-file lease, where the writer marks it, and fails with the same "volume is unavailable: <reason>" error the old check returned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: read a non-ASCII or empty Range header as Go does A Range value with a byte >= 0x80 (obs-text, which hyper accepts) failed HeaderValue::to_str at both range gates. For a needle served as stored the handler had already chosen a meta-only read, so it fell through to the buffered path with no payload and answered 200 with an empty body; a compressed or EC needle answered 200 with the full body. Go's parseRange fails on a byte it can neither trim nor parse and answers 416 "invalid range", and trims Unicode whitespace such as NBSP into a normal 206. An empty Range value was also a 200 with an empty body, where Go sends the whole payload. Read Range once with from_utf8_lossy, dropping an empty value, and hand that one value to the read plan and to both range gates. A replaced byte never parses, so it is a 416; str::trim trims the same Unicode whitespace as strings.TrimSpace. A range read from the data file now always answers itself instead of falling through with an empty needle. The buffered path answers HEAD before it looks at Range, as Go's writeResponseContent does, so an EC HEAD with a Range is a 200 with the full length. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
16e66b1bad |
fix(s3): initialize destination ACLs for CopyObject (#11599)
* fix(s3): initialize destination ACLs for CopyObject Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * fix(s3): re-check routed self-copy eligibility on the locked read routeInPlace was decided on the pre-lock entry, but the PATCH body re-reads the entry. A concurrent write changing file mode or MIME in between left the routed PATCH installing new ACL keys while Attributes kept the stale mode. Evaluate eligibility against the re-read entry and retry the self-copy under the distributed lock when it no longer qualifies. * fix(s3): guard metadata self-copies against concurrent writes Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * ci: raise s3api unit-test timeout to 9m The suite crossed the 5m binary timeout on the hosted runner (local run is ~4.3m and still growing). The job-level limit is already 10m. * ci: allow setup time before the S3 API test suite Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> --------- Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
7e809c9991 |
admin: stop leaking cancelled maintenance task files in -dataDir/tasks (#11597)
* admin: delete persisted state when scan cancels pending tasks Each detection cycle cancels every pending task of a type before re-detecting it, and the cancel path saved the cancelled task back to disk. Nothing ever removed those files, so -dataDir/tasks gained one orphaned .pb per candidate volume per scan cycle. Cancelled is terminal, so drop the file the same way CompleteTask does for completed/failed tasks. The cancelled entry stays in memory for the UI until the next purge. Refs #11595 * admin: delete persisted state when CancelTask cancels a pending task The manual cancel path only updated memory, leaving the pending .pb on disk where a restart would resurrect the cancelled task as pending and the file would linger until then. Delete it like the scan-cycle cancel path now does. * admin: count cancelled tasks toward task retention cleanup CleanupOldTasks and ConfigPersistence.CleanupCompletedTasks only filtered completed/failed tasks, so cancelled entries were exempt from retention in both memory and on disk. Treat all terminal states alike; nil CompletedAt entries also count and sort last, so they are pruned first. * admin: run task file retention in the periodic cleanup loop cleanupCompletedTasks had no callers, so the on-disk retention bound never ran during uptime. Invoke it from performCleanup alongside the in-memory CleanupOldTasks sweep. * admin: guard task state writes against stale saves and failed deletes saveTaskState runs after mq.mutex is released, so the task may have gone terminal in between; a delayed pending save could then recreate the file a cancel just deleted and resurrect the task on restart. Skip saving non-terminal snapshots once the live task is terminal or gone. If a cancel file removal fails, fall back to writing the cancelled snapshot so the file is terminal rather than pending. deleteTaskState now returns its error, and CancelTask captures task.Status while still holding the queue lock. * admin: serialize task file check+write against cancel deletes The saveTaskState guard still had a check-then-write window: a pending snapshot could pass the terminal check before a cancel deleted the file, then write it back after. A persistMu on the queue now covers the check+save and the cancel paths' delete (with its terminal-state fallback), so the two cannot interleave for the same task. |
||
|
|
3c17c5146e |
S3: fix ListObjectVersions losing keys across page boundaries (#11598)
* s3api: thread filer client through the versioned-listing collector
findVersionsRecursively now binds one SeaweedFilerClient for the whole
recursive walk instead of re-resolving a filer on every list/lookup call,
and the collector's list/getEntry/scanLatestVersionEntry/getObjectVersionList
helpers go through it. No behavior change; this also lets tests drive
collectVersions with a stubbed client.
* s3api: keep collecting versions while pending names can sort into the page
ListObjectVersions walked the filer in directory-entry name order and
stopped as soon as maxKeys+1 items were collected, sorting only that
partial set. Filer names do not match key order: "a.copy.versions" sorts
before "a.versions" while key "a.copy" sorts after "a", so a page
boundary inside the earlier-walked sibling's versions permanently skipped
the later key.
Track the largest key collected (maxKey) and, once the collector is full,
keep walking until entry names pass the ceiling of names that can still
resolve to keys at or below it; the ceiling reaches through the prefix
versions of maxKey. Entries whose subtree can only hold keys above maxKey
are skipped. Versions of an in-bound object are collected in full so its
position in the sorted page is exact.
Fixes seaweedfs#11594
* s3api: resume versioned listings at the earliest covering name prefix
computeStartFrom mapped the key marker straight to an entry name (or cut
it at the first '/'), which skips sibling directories that are a prefix
of the marker below '0' - for marker "d.x" the listing resumed at name
"d.x", skipping directory "d" whose keys "d/*" all sort after it.
Resume at the earliest remainder prefix ending at a byte below '0' ('/',
'.', '-' and friends), so every directory whose subtree can still hold
keys past the marker is revisited; already-returned keys inside are
filtered by the existing marker checks as before.
* s3api: regression test for versioned-listing pagination order
Drive collectVersions against a stubbed filer holding the issue-11594
layout - "a.copy.versions" listing before "a.versions", plus a "d/"
subtree next to "d.x" - and assert that every page size from 1 up
reproduces the unpaginated ordering with no lost or duplicated entries.
Also updates TestComputeStartFrom for the new earliest-prefix resume and
gives testFilerClient a LookupDirectoryEntry stub.
* s3api: inject list/getEntry functions into the version collector
Pinning one SeaweedFilerClient for the whole walk dropped per-call
failover: previously each s3a.list resolved a filer through
WithFilerClient, so a mid-walk filer failure could fall back to a
healthy peer. Inject s3a.list/s3a.getEntry as function fields instead -
production keeps the failover behavior, tests can still stub.
* s3api: keep scanning marker for a later covering prefix
A leading byte below '0' (marker .hidden/file) has no non-empty prefix
at index 0, but a deeper separator still does - resuming at .hidden/file
skipped the .hidden directory and its keys after file. Continue the scan
instead of bailing on the first byte.
|
||
|
|
483dd4b12e |
s3api: persist ACLs on PutObject uploads (#11592)
* s3api: persist ACLs on PutObject uploads Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * s3api: fix PutObject ACL edge cases found in review - Only enforce BucketOwnerEnforced when explicitly configured; buckets without a stored ownership control keep accepting upload ACLs - Ignore ACL query parameters on SigV2 requests, which do not sign them - Mirror signed-query ACL values into headers after authentication so grant parsing and resolveFileMode agree on presigned uploads - Validate only caller-supplied grantees against the account registry; default grants now work for accounts outside the local registry - Reject unknown grantee keys and accept comma-separated grantee lists without spaces in ParseCustomAclHeader - Guard against identities without an account * s3api: harden upload ACL parsing and authorization Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * s3api: evaluate upload ACL grantees individually in policies A comma-joined grant header or a signed query parameter reached policy conditions as one value, so a deny on a later grantee did not fire. Split grant headers into per-grantee values for policy evaluation and share the grantee pair parser with ParseCustomAclHeader. * s3api: keep raw grant header values visible to policy conditions Exact-match conditions written against the signed header value stopped matching once grantees were split for evaluation. Preserve the original wire values alongside the per-grantee values so deny policies fire on either granularity. * s3api: evaluate upload ACL grants as one canonical list in policies Conditions on s3:x-amz-grant-* now see a single comma-separated canonical grant list identical for a single line, repeated header lines, or a signed query parameter. This keeps StringEquals allows and exact-list or allowlist (StringNotEquals) denies accurate regardless of wire encoding. * s3api: preserve upload ACL denies and align policy checks Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * s3api: retain upload owner grants and literal policy values Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> --------- Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
9d1c24d80d |
[Filer] Support append to inline small files (#11591)
* fix 11586 * Update filer_server_handlers_write_autochunk.go * filer: fix inline append races, empty files, and stale ETags Serialize the append read-modify-write on the entry lock so concurrent appends merge instead of losing content, keep small appends to empty files inline, tolerate legacy entries whose metadata size differs from their content, and set the entry digest so appended inline files keep a real ETag. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
76ddde6a6d | docs: regenerate star history chart | ||
|
|
0d93dec145 | helm: mount TLS certificates in bucket hook (#11589) | ||
|
|
d9b69a7f76 |
s3api: fix PutObjectAcl permission scoping and owner grants (#11587)
PutObjectAcl had four authorization and ownership bugs: - The handler embedded the resource path into the action (WriteAcp:bucket/object), and authRequest/CanDo then scoped it to the request's bucket/object again. A bucket-wide WriteAcp:bucket grant could never match, so legitimate owners got 403. - After authRequest succeeded via an IAM or bucket policy, a leftover identity.CanDo gate re-checked only the legacy Actions list, denying identities authorized purely by policies. - For canned and default ACLs, ExtractAcl generated the FULL_CONTROL grant for the requesting account instead of the object owner. An admin setting private/public-read on another account's object left the owner metadata intact but reassigned full control to the admin. - Objects without stored owner metadata (e.g. written via the filer outside S3) fell back to treating the requester as the owner, so any user with a WriteAcp grant could take them over. Non-admins are now denied; admins keep the takeover fallback. Grantee validation now also accepts the object's stored owner even when that account has been removed from the registry, so canned/XML ACLs for retired owners keep working. |
||
|
|
2b5fdc639f |
filer: stop isSameChunks from sorting caller-owned chunk slices (#11584)
* filer: stop isSameChunks from sorting caller-owned chunk slices slices.SortFunc reorders the input in place. filer.remote.sync calls IsSameData on a metadata event's NewEntry inside isMetadataOnlyUpdate and later stamps the filer entry under an IF_ENTRY_EQUAL precondition carrying that same entry. The ETag-sorted chunk list never matches the stored entry, so every stamp of a multi-chunk object fails, synced_mtime_ns stays zero, and dirty objects are re-uploaded forever. Sort clones of the slices instead. * filer: test IsSameData leaves input chunk order unchanged Guards the clone-then-sort fix: a regression back to in-place sorting would reorder caller-owned chunk slices and reintroduce the remote-sync IF_ENTRY_EQUAL mismatch. |
||
|
|
d7a02567e3 | docs: regenerate star history chart | ||
|
|
39bc9cd0ef |
s3api: copy the trailer checksum before reading the next trailer line (#11583)
* s3api: copy the trailer checksum before reading the next trailer line
parseChunkChecksum kept the checksum value as a sub-slice of the line
returned by bufio.Reader.ReadSlice, which is only valid until the next
read. When the trailer lines arrive in separate TCP segments, reading
x-amz-trailer-signature refills the buffer and overwrites the saved
value, so a correct upload fails with InvalidDigest ("The Content-Md5
you specified is not valid").
The AWS SDK for Java v2 (>= 2.30) on a Linux JDK sends the trailer that
way; about half of its signed streaming uploads failed.
Fixes #11582
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* s3api: reuse crc32 writer and trim comments in trailer split test
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
|
||
|
|
f4ef37e752 |
filer.sync: resubscribe the metadata stream when a failure pins the offset (#11581)
* filer sink: keep the gRPC status inside wrapped errors
%v stringifies the status, so a peer teardown reported as Canceled ("the
client connection is closing") reached IsTransientError as plain text and
matched nothing: the sync job failed on the first attempt and pinned the
offset. %w keeps the status reachable, so the retry runs on a fresh
connection once the target is back.
* pb: let a consumer drop the metadata stream to force a resubscribe
A MetadataProcessor job that exhausts its retries pins the processed
watermark so the event replays on the next subscribe — but nothing on the
source stream notices a target-side failure, so the replay waited for an
unrelated reconnect or a restart. The new Resubscribe channel cancels the
stream's context; the Recv loop answers it with ErrResubscribe so the
caller's retry loop resubscribes from GetResumeTsNs and replays the pinned
events in order.
* pb: stop the event retry loop once the stream context is done
RetryUntil ignores context, so a subscriber parked on a failing offset
write would keep retrying past a resubscribe signal until the sink came
back. Stop retrying when the stream is being dropped so the resubscribe
takes effect promptly.
* filer.sync: signal resubscribe when a job failure pins the offset
A job that exhausts its in-job retries leaves the event pinned behind oldestFailedTsNs, replayable only on a reconnect. Closing resubscribeCh on the first recorded failure lets the metadata follower drop the stream so the reconnect replays the pinned events instead of waiting for a process restart (#11572).
* filer.sync: wire the resubscribe signal into the follow options
filer.sync, filer.remote.sync, and the remote gateway bucket sync all run their subscription inside an outer retry loop, so ErrResubscribe resurfaces as a resubscribe from the persisted watermark.
* filer.sync: wait for in-flight jobs before signaling resubscribe
* remote sync: never resume past the saved offset when -timeAgo is set
* filer.sync: drop events that arrive after the drain signals resubscribe
* pb: interrupt the event retry backoff when the stream context ends
* filer.sync: stop admitting once a failure pins, and count jobs per timestamp
A pinned watermark only released once the processor went fully quiet, so a busy stream could starve the resubscribe — the failed event would wait for an unrelated reconnect anyway, the wait this mechanism exists to remove. The processor now latches stopped when a job fails: admission drops new events (they replay from the pinned watermark after the reconnect), a broadcast releases blocked waiters, and the resubscribe signals as soon as the jobs already in flight drain. A redelivery of an event still in the failure ledger may still run so its success shrinks the replay, but nothing starts once the signal has fired, or it would race the replay it asked for.
Dropped events no longer inflate the received counters — an event counts only once admitted, and the replay's own admission counts it.
While here: activeJobs keyed by TsNs collapsed events sharing a timestamp, so one completion could empty the map while a same-ts sibling was still running — letting the drain gate and the watermark outrun it. Jobs are now counted per timestamp, and the drain and lazy heap cleanup go through the counts.
|
||
|
|
10b0f2b8ad |
volume server: refuse the rest of a grouped run after a durable index failure (#11576)
* volume server: refuse the rest of a grouped run after a durable index failure A durable write whose needle-map put fails stops the volume taking writes (#10825): sent on its own, the next write then fails read only before it appends. The grouped run from #11543 appends and syncs every entry before publishing any, then kept publishing the entries after the failed one and acked them once the shared .idx sync went through. When the failed put tore its .idx row, the rows appended after it land off alignment, so the next load parses them as garbage and the acked writes are gone. Once a durable entry fails to publish, refuse every later entry of the run with ReadOnly, as the per-needle path does. The entries before it stay acked; their rows go down with the run's one .idx sync. The refused records stay on the .dat unindexed, as the failed one does on its own. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: refuse a grouped entry staged as a cookie mismatch too After a durable entry in a grouped run fails to index, the entries after it are refused as they would be on their own. On its own an entry meets check_writable before its cookie check, so one staged as a cookie mismatch now gets the refusal too, instead of keeping its staging error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: trim a torn .idx row back so the next stays aligned A failed write_index_entry can leave half a row in the .idx. With the writer appending at the tail, every row written after it lands off alignment and the next load parses them as garbage, so a write acked behind a torn row does not come back. Trim the file back to idx_file_offset on a failed append, in both needle maps, and cover it with a test that writes past a torn row and reloads. * volume server: refuse queued Go writes once a durable index update fails processBatch kept writing after a failed nm.Put, and the single-write path checked IsReadOnly only outside the volume lock. A durable write whose index update fails now marks the volume noWriteOrDelete, and each queued request is checked before it appends, so the ones after a failed durable entry are refused the way a lone write is. Deletes get the same noWriteOrDelete refusal a lone delete gets. * volume server: refuse appends while a torn .idx row cannot be trimmed When trimming back a half-written .idx row itself fails, the next append would land after the torn bytes and every later row would parse off alignment on load. Latch the map as torn and refuse appends until the trim succeeds, on both CompactNeedleMap and RedbNeedleMap; the same latch covers an orphan row that could not be trimmed after a failed redb commit. The .idx writer is now opened with write+append access so truncate_to (set_len) works on Windows, where an append-only handle cannot trim. * volume server: write .idx rows at idx_file_offset, not via append mode Rust's OpenOptions on Windows strips FILE_WRITE_DATA whenever append is set so the handle stays strictly append-only, which makes set_len fail - the torn-row trim could never succeed there. Open the .idx writer with plain write access and seek to idx_file_offset before each row, the same positioned-write model the Go server uses. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
eafe79ebff |
filer: skip UpdateEntry when inline content is unchanged (#11580)
* filer: skip UpdateEntry when inline content is unchanged SaveInsideFiler rewrites config files (IAM identities, filer.conf, remote mappings, policies) unconditionally. Each no-op UpdateEntry is a metadata event the local meta log persists to /topics/.system/log, which appends a chunk to a volume. A client that rewrites identical config on a timer, e.g. the seaweedfs-operator 5-minute resync calling UpdateUser with unchanged actions, keeps .dat/.idx files growing on an otherwise idle cluster and prevents HDD spindown (seaweedfs/seaweedfs#11571). Skip the UpdateEntry when the stored inline content is byte-identical, so unchanged writes produce no metadata event and no volume writes. * filer: test that identical SaveInsideFiler writes skip UpdateEntry * filer: require stamped Md5 before skipping identical writes An entry holding identical content but no Md5 (written before hashing, or by a tool that cleared it) would never get the stamp that IF_ETAG_MATCH conditional writes key off. Skip only when both the stored content and its Md5 match, so one write still lands to repair the stamp. |