mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-06 14:31:57 +02:00
52c6df3beca29c5f5716bf23b786663e73fb1b60
15390
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
52c6df3bec |
fix(filer): leave the lock ring before stopping gRPC on shutdown (#11604)
A draining filer stayed in the master's lock ring until its process exited, while gRPC GracefulStop was already refusing new connections. S3 gateways and peer filers kept routing object-write locks and owner-routed writes to it for the whole graceful-stop window. On shutdown the filer now sends leave_lock_ring on its open KeepConnected stream. The master removes it from the lock ring only; it stays a cluster member so peers keep following its metadata log through the drain. Once the ring update without it arrives, the filer has already transferred its locks to the new owners, and it keeps serving through the prior-owner window (plus a second for peers that apply the update later) before gRPC and HTTP begin draining. The whole leave is bounded at 10s so slow lock transfers or a stuck stream cannot hold up the drain. A lone filer, or a master that ignores the message, falls back to the previous behavior. |
||
|
|
aa5b337716 |
fix(s3): honor object ACLs for anonymous GET and HEAD (#11605)
* fix(s3): honor object ACLs for anonymous reads Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * s3: drop unsigned session tokens on deferred anonymous reads An unsigned GET/HEAD carrying X-Amz-Security-Token is anonymous, not a session request; strip the token before authentication and authorization so it cannot influence identity or policy evaluation. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
abdb95dc87 |
build(deps): bump golang.org/x/time from 0.15.0 to 0.16.0 (#11608)
Bumps [golang.org/x/time](https://github.com/golang/time) from 0.15.0 to 0.16.0. - [Commits](https://github.com/golang/time/compare/v0.15.0...v0.16.0) --- updated-dependencies: - dependency-name: golang.org/x/time dependency-version: 0.16.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
076fd24186 |
filer: bound metadata log flush retries during shutdown (#11596)
* filer: bound metadata log flush retries during shutdown On SIGTERM the filer could hang forever in Shutdown: the final LocalMetaLogBuffer flush retries appendToFile indefinitely, and with the master already down each AssignVolume attempt just kept failing. WaitForShutdown never returned, the interrupt hook never reached os.Exit, and the half-dead filer kept its ports bound. Thread a context through appendToFile/assignAndUpload and switch to a 15s-bounded context once the filer is stopping: the flush abandons with a log line instead of retrying forever. Normal operation keeps the unbounded retry so no metadata is dropped while running. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: share one shutdown deadline across all pending meta log flushes Review feedback on the per-flush timeout: a deadline armed at flush start could already be expired when shutdown arrived, the first append of a flush still ran unbounded, each queued window got a fresh budget (16 windows * 15s), and an abandoned window still advanced the flushed watermark as if it had landed. Rework to a single shared flush context on the Filer, cancelled once by Shutdown via AfterFunc. Every append - the in-flight one and every queued window - observes the same deadline, so the whole drain is bounded at 15s. A flush that gives up reports its dropped bytes through the new LogBuffer.NoteFlushDropped, and loopFlush then skips the offset/timestamp advance and subscriber notifications for that window. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: bound append attempts and commit uploaded pieces detached Store operations check then drop request cancellation, so a stalled backend could still hold flushFn past the shutdown deadline; run each append attempt on its own goroutine and give up on it at the deadline. Once a piece is uploaded, commit its entry on a detached context so the expired deadline cannot strand the chunk. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: keep shutdown flushes synchronous Detaching the append attempt let flushFn return while the goroutine still held the pooled flush buffer and could commit after the metadata store closed; abandonment is only safe for the cancelable assign/upload phase, which the shared flush context already bounds. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
90eb4ec091 |
build(deps): bump github.com/hashicorp/raft-boltdb/v2 from 2.3.1 to 2.4.2 (#11606)
build(deps): bump github.com/hashicorp/raft-boltdb/v2 Bumps [github.com/hashicorp/raft-boltdb/v2](https://github.com/hashicorp/raft-boltdb) from 2.3.1 to 2.4.2. - [Release notes](https://github.com/hashicorp/raft-boltdb/releases) - [Commits](https://github.com/hashicorp/raft-boltdb/compare/v2.3.1...v2.4.2) --- updated-dependencies: - dependency-name: github.com/hashicorp/raft-boltdb/v2 dependency-version: 2.4.2 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
e710cfc4b0 |
build(deps): bump github.com/rabbitmq/amqp091-go from 1.14.0 to 1.15.0 (#11610)
Bumps [github.com/rabbitmq/amqp091-go](https://github.com/rabbitmq/amqp091-go) from 1.14.0 to 1.15.0. - [Release notes](https://github.com/rabbitmq/amqp091-go/releases) - [Changelog](https://github.com/rabbitmq/amqp091-go/blob/main/CHANGELOG.md) - [Commits](https://github.com/rabbitmq/amqp091-go/compare/v1.14.0...v1.15.0) --- updated-dependencies: - dependency-name: github.com/rabbitmq/amqp091-go dependency-version: 1.15.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
cf144d5eb2 | docs: regenerate star history chart | ||
|
|
825c3dff8b |
volume server: fetch only the chunks a manifest Range needs, like Go (#11546)
* volume server: read GET/HEAD needles off the store lock, and only once The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stream needle chunks without the store lock StreamingBody::poll_frame took store.read() and find_volume for every chunk to compare the volume's compaction revision, dup'd the source handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the stream also holds a data-file read lease for its whole life, while a writer waits for that lease under store.write(): the next chunk's store.read() then waits for the writer and the writer for the stream. The per-chunk re-lookup was also wrong. The stream reads a handle opened at plan time, which pins the .dat inode the offset was resolved against; a vacuum commit renames a new file over .dat and leaves that inode untouched. The re-looked-up offset belongs to the new file but was read from the old inode, so a stream whose needle a vacuum moved ended in a checksum error. The pinned offset stays valid, so the check, and with it every store access, is dropped, along with the now unused re_lookup_needle_data_offset and the revision fields of the read plan. The source is shared as an Arc instead of dup'd per chunk, and the chunk buffer is a BytesMut that the blocking read hands back with its result, so its allocation is reclaimed once the previous frame has been written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: split get_or_head_handler_inner into phases get_or_head_handler_inner was a ~650-line function. Its middle resolved the needle and set five mutable flags (stream_info, can_stream, can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that three if-let reply paths then re-tested, each re-checking stream_info. It is now a 126-line orchestrator over named phases: reject_read_jwt, proxy_missing_volume, wait_for_download_slot, parse_read_request, read_ec_needle / read_volume_needle, etag_and_last_modified, not_modified_response, read_response_headers, and the reply phases stream_response, head_from_meta_response, range_from_source_response, buffered_payload and buffered_response. The read phases return a ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource, Buffered) carries the NeedleStreamInfo only on the variants that use it, so the reply is one match instead of three flag checks. Pure refactor: every status code, header and header order, error text, metric increment, lock and data-file lease scope, spawn_blocking boundary and side-effect order is unchanged. Phases that can end the request return ControlFlow<Response, T>. A Range header that is not visible ASCII still falls through to the buffered path, as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a needle stream once its volume becomes unavailable Taking the store lock out of StreamingBody also dropped its per-chunk unavailable_error() check. With -hasSlowRead a writer can take the data-file lease between chunks, fail its fsync and its truncate, and mark the volume unavailable; the stream then kept serving the rest of the needle from its pinned handle. The volume's io_unavailable reason is now an Arc-shared leaf mutex that the read plan hands to the stream. Each chunk checks it under its data-file lease, where the writer marks it, and fails with the same "volume is unavailable: <reason>" error the old check returned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: read a non-ASCII or empty Range header as Go does A Range value with a byte >= 0x80 (obs-text, which hyper accepts) failed HeaderValue::to_str at both range gates. For a needle served as stored the handler had already chosen a meta-only read, so it fell through to the buffered path with no payload and answered 200 with an empty body; a compressed or EC needle answered 200 with the full body. Go's parseRange fails on a byte it can neither trim nor parse and answers 416 "invalid range", and trims Unicode whitespace such as NBSP into a normal 206. An empty Range value was also a 200 with an empty body, where Go sends the whole payload. Read Range once with from_utf8_lossy, dropping an empty value, and hand that one value to the read plan and to both range gates. A replaced byte never parses, so it is a 416; str::trim trims the same Unicode whitespace as strings.TrimSpace. A range read from the data file now always answers itself instead of falling through with an empty needle. The buffered path answers HEAD before it looks at Range, as Go's writeResponseContent does, so an EC HEAD with a Range is a 200 with the full length. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: apply Range to chunk manifests and forward raw headers when proxying A GET of a chunk manifest assembled the object and always answered 200 with the whole body, ignoring Range. Go serves the expanded manifest through writeResponseContent, which answers HEAD first and then hands Range to ProcessRangeRequest: 206 for one range, multipart/byteranges for several, 416 for an unsatisfiable or unparsable one. try_expand_chunk_manifest now returns the assembled body and headers, and the caller answers through buffered_response, the same helper the buffered needle path uses. A proxied read forwarded a request header only if HeaderValue::to_str succeeded, so a Range with an obs-text byte was dropped and the target answered 200 with the full body. Go copies every header value as is. Forward the raw HeaderValue for every header. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: fetch only the chunks a manifest Range needs, like Go A ranged GET of a chunk manifest fetched every chunk, assembled the whole object and then sliced it, so reading a few bytes of a large object cost a read of all of it, and a 416 still fetched everything. Go serves a manifest through ChunkedFileReader, which seeks to each range and reads only the chunks under it. For a GET with a Range, try_expand_chunk_manifest now parses the ranges against the manifest size, fetches only the chunks whose declared window overlaps one of them (none when the reply carries no body), and answers through handle_range_request_with, the reader-based core that handle_range_request now wraps, so 206/416/multipart stay one code path. The reader replays assembly: chunks clamped as before, later chunks over earlier ones, zeros in gaps. HEAD, no-Range GETs and GETs that crop or resize an image still assemble the whole object. A missing chunk outside the requested ranges no longer turns a ranged GET into a 500, as in Go; a missing chunk inside them still does, before any headers are sent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: reword a comment codespell flags Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: keep only range-covered bytes of fetched manifest chunks A ranged GET retained every overlapping chunk's full contents; 1,000 overlapping 8 MiB chunks could pin ~8 GiB for a one-byte response. Clip each fetched chunk to the bytes the requested ranges can actually read, preserving the later-chunks-overwrite and zero-fill-gap semantics. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: bucket ranged manifest parts by range Serving a multipart range scanned every retained part. Bucket the kept intersections by their range so one range only reads its own parts. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
a7a590474d |
kafka: honor notification.kafka.event_types (#11601)
* kafka: honor notification.kafka.event_types Kafka published every filer event and ignored the filter the webhook notifier already uses. Co-authored-by: Cursor <cursoragent@cursor.com> * notification: share event-type classification between queues Kafka duplicated the webhook's event classification verbatim; move it to the notification package so the two queues cannot drift. Webhook keeps its typed eventType wrappers over the shared helpers. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
8c67b75190 |
image: add an optional public image processing gateway (#11593)
* image: add an optional public image processing gateway * image: fix representation metadata and processing bounds * image: restrict passthrough to non-executable media types Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * image: tighten source media-type validation Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * image: write passthrough body on the inner response writer CodeQL still flagged the passthrough write: the content type was set on the wrapper while the body reached w.ResponseWriter, so the validated header could not be associated with the write. Set headers and copy the body on the same inner writer. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * image: serve processed output on the inner response writer * image: reject XML source types and unsafe conditional metadata --------- Co-authored-by: zhaoyuchen <yc.zhao@yinzon.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
7a19961074 |
volume server: read a non-ASCII or empty Range header as Go does (#11534)
* volume server: read GET/HEAD needles off the store lock, and only once The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stream needle chunks without the store lock StreamingBody::poll_frame took store.read() and find_volume for every chunk to compare the volume's compaction revision, dup'd the source handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the stream also holds a data-file read lease for its whole life, while a writer waits for that lease under store.write(): the next chunk's store.read() then waits for the writer and the writer for the stream. The per-chunk re-lookup was also wrong. The stream reads a handle opened at plan time, which pins the .dat inode the offset was resolved against; a vacuum commit renames a new file over .dat and leaves that inode untouched. The re-looked-up offset belongs to the new file but was read from the old inode, so a stream whose needle a vacuum moved ended in a checksum error. The pinned offset stays valid, so the check, and with it every store access, is dropped, along with the now unused re_lookup_needle_data_offset and the revision fields of the read plan. The source is shared as an Arc instead of dup'd per chunk, and the chunk buffer is a BytesMut that the blocking read hands back with its result, so its allocation is reclaimed once the previous frame has been written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: split get_or_head_handler_inner into phases get_or_head_handler_inner was a ~650-line function. Its middle resolved the needle and set five mutable flags (stream_info, can_stream, can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that three if-let reply paths then re-tested, each re-checking stream_info. It is now a 126-line orchestrator over named phases: reject_read_jwt, proxy_missing_volume, wait_for_download_slot, parse_read_request, read_ec_needle / read_volume_needle, etag_and_last_modified, not_modified_response, read_response_headers, and the reply phases stream_response, head_from_meta_response, range_from_source_response, buffered_payload and buffered_response. The read phases return a ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource, Buffered) carries the NeedleStreamInfo only on the variants that use it, so the reply is one match instead of three flag checks. Pure refactor: every status code, header and header order, error text, metric increment, lock and data-file lease scope, spawn_blocking boundary and side-effect order is unchanged. Phases that can end the request return ControlFlow<Response, T>. A Range header that is not visible ASCII still falls through to the buffered path, as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a needle stream once its volume becomes unavailable Taking the store lock out of StreamingBody also dropped its per-chunk unavailable_error() check. With -hasSlowRead a writer can take the data-file lease between chunks, fail its fsync and its truncate, and mark the volume unavailable; the stream then kept serving the rest of the needle from its pinned handle. The volume's io_unavailable reason is now an Arc-shared leaf mutex that the read plan hands to the stream. Each chunk checks it under its data-file lease, where the writer marks it, and fails with the same "volume is unavailable: <reason>" error the old check returned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: read a non-ASCII or empty Range header as Go does A Range value with a byte >= 0x80 (obs-text, which hyper accepts) failed HeaderValue::to_str at both range gates. For a needle served as stored the handler had already chosen a meta-only read, so it fell through to the buffered path with no payload and answered 200 with an empty body; a compressed or EC needle answered 200 with the full body. Go's parseRange fails on a byte it can neither trim nor parse and answers 416 "invalid range", and trims Unicode whitespace such as NBSP into a normal 206. An empty Range value was also a 200 with an empty body, where Go sends the whole payload. Read Range once with from_utf8_lossy, dropping an empty value, and hand that one value to the read plan and to both range gates. A replaced byte never parses, so it is a 416; str::trim trims the same Unicode whitespace as strings.TrimSpace. A range read from the data file now always answers itself instead of falling through with an empty needle. The buffered path answers HEAD before it looks at Range, as Go's writeResponseContent does, so an EC HEAD with a Range is a 200 with the full length. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
16e66b1bad |
fix(s3): initialize destination ACLs for CopyObject (#11599)
* fix(s3): initialize destination ACLs for CopyObject Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * fix(s3): re-check routed self-copy eligibility on the locked read routeInPlace was decided on the pre-lock entry, but the PATCH body re-reads the entry. A concurrent write changing file mode or MIME in between left the routed PATCH installing new ACL keys while Attributes kept the stale mode. Evaluate eligibility against the re-read entry and retry the self-copy under the distributed lock when it no longer qualifies. * fix(s3): guard metadata self-copies against concurrent writes Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * ci: raise s3api unit-test timeout to 9m The suite crossed the 5m binary timeout on the hosted runner (local run is ~4.3m and still growing). The job-level limit is already 10m. * ci: allow setup time before the S3 API test suite Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> --------- Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
7e809c9991 |
admin: stop leaking cancelled maintenance task files in -dataDir/tasks (#11597)
* admin: delete persisted state when scan cancels pending tasks Each detection cycle cancels every pending task of a type before re-detecting it, and the cancel path saved the cancelled task back to disk. Nothing ever removed those files, so -dataDir/tasks gained one orphaned .pb per candidate volume per scan cycle. Cancelled is terminal, so drop the file the same way CompleteTask does for completed/failed tasks. The cancelled entry stays in memory for the UI until the next purge. Refs #11595 * admin: delete persisted state when CancelTask cancels a pending task The manual cancel path only updated memory, leaving the pending .pb on disk where a restart would resurrect the cancelled task as pending and the file would linger until then. Delete it like the scan-cycle cancel path now does. * admin: count cancelled tasks toward task retention cleanup CleanupOldTasks and ConfigPersistence.CleanupCompletedTasks only filtered completed/failed tasks, so cancelled entries were exempt from retention in both memory and on disk. Treat all terminal states alike; nil CompletedAt entries also count and sort last, so they are pruned first. * admin: run task file retention in the periodic cleanup loop cleanupCompletedTasks had no callers, so the on-disk retention bound never ran during uptime. Invoke it from performCleanup alongside the in-memory CleanupOldTasks sweep. * admin: guard task state writes against stale saves and failed deletes saveTaskState runs after mq.mutex is released, so the task may have gone terminal in between; a delayed pending save could then recreate the file a cancel just deleted and resurrect the task on restart. Skip saving non-terminal snapshots once the live task is terminal or gone. If a cancel file removal fails, fall back to writing the cancelled snapshot so the file is terminal rather than pending. deleteTaskState now returns its error, and CancelTask captures task.Status while still holding the queue lock. * admin: serialize task file check+write against cancel deletes The saveTaskState guard still had a check-then-write window: a pending snapshot could pass the terminal check before a cancel deleted the file, then write it back after. A persistMu on the queue now covers the check+save and the cancel paths' delete (with its terminal-state fallback), so the two cannot interleave for the same task. |
||
|
|
3c17c5146e |
S3: fix ListObjectVersions losing keys across page boundaries (#11598)
* s3api: thread filer client through the versioned-listing collector
findVersionsRecursively now binds one SeaweedFilerClient for the whole
recursive walk instead of re-resolving a filer on every list/lookup call,
and the collector's list/getEntry/scanLatestVersionEntry/getObjectVersionList
helpers go through it. No behavior change; this also lets tests drive
collectVersions with a stubbed client.
* s3api: keep collecting versions while pending names can sort into the page
ListObjectVersions walked the filer in directory-entry name order and
stopped as soon as maxKeys+1 items were collected, sorting only that
partial set. Filer names do not match key order: "a.copy.versions" sorts
before "a.versions" while key "a.copy" sorts after "a", so a page
boundary inside the earlier-walked sibling's versions permanently skipped
the later key.
Track the largest key collected (maxKey) and, once the collector is full,
keep walking until entry names pass the ceiling of names that can still
resolve to keys at or below it; the ceiling reaches through the prefix
versions of maxKey. Entries whose subtree can only hold keys above maxKey
are skipped. Versions of an in-bound object are collected in full so its
position in the sorted page is exact.
Fixes seaweedfs#11594
* s3api: resume versioned listings at the earliest covering name prefix
computeStartFrom mapped the key marker straight to an entry name (or cut
it at the first '/'), which skips sibling directories that are a prefix
of the marker below '0' - for marker "d.x" the listing resumed at name
"d.x", skipping directory "d" whose keys "d/*" all sort after it.
Resume at the earliest remainder prefix ending at a byte below '0' ('/',
'.', '-' and friends), so every directory whose subtree can still hold
keys past the marker is revisited; already-returned keys inside are
filtered by the existing marker checks as before.
* s3api: regression test for versioned-listing pagination order
Drive collectVersions against a stubbed filer holding the issue-11594
layout - "a.copy.versions" listing before "a.versions", plus a "d/"
subtree next to "d.x" - and assert that every page size from 1 up
reproduces the unpaginated ordering with no lost or duplicated entries.
Also updates TestComputeStartFrom for the new earliest-prefix resume and
gives testFilerClient a LookupDirectoryEntry stub.
* s3api: inject list/getEntry functions into the version collector
Pinning one SeaweedFilerClient for the whole walk dropped per-call
failover: previously each s3a.list resolved a filer through
WithFilerClient, so a mid-walk filer failure could fall back to a
healthy peer. Inject s3a.list/s3a.getEntry as function fields instead -
production keeps the failover behavior, tests can still stub.
* s3api: keep scanning marker for a later covering prefix
A leading byte below '0' (marker .hidden/file) has no non-empty prefix
at index 0, but a deeper separator still does - resuming at .hidden/file
skipped the .hidden directory and its keys after file. Continue the scan
instead of bailing on the first byte.
|
||
|
|
483dd4b12e |
s3api: persist ACLs on PutObject uploads (#11592)
* s3api: persist ACLs on PutObject uploads Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * s3api: fix PutObject ACL edge cases found in review - Only enforce BucketOwnerEnforced when explicitly configured; buckets without a stored ownership control keep accepting upload ACLs - Ignore ACL query parameters on SigV2 requests, which do not sign them - Mirror signed-query ACL values into headers after authentication so grant parsing and resolveFileMode agree on presigned uploads - Validate only caller-supplied grantees against the account registry; default grants now work for accounts outside the local registry - Reject unknown grantee keys and accept comma-separated grantee lists without spaces in ParseCustomAclHeader - Guard against identities without an account * s3api: harden upload ACL parsing and authorization Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * s3api: evaluate upload ACL grantees individually in policies A comma-joined grant header or a signed query parameter reached policy conditions as one value, so a deny on a later grantee did not fire. Split grant headers into per-grantee values for policy evaluation and share the grantee pair parser with ParseCustomAclHeader. * s3api: keep raw grant header values visible to policy conditions Exact-match conditions written against the signed header value stopped matching once grantees were split for evaluation. Preserve the original wire values alongside the per-grantee values so deny policies fire on either granularity. * s3api: evaluate upload ACL grants as one canonical list in policies Conditions on s3:x-amz-grant-* now see a single comma-separated canonical grant list identical for a single line, repeated header lines, or a signed query parameter. This keeps StringEquals allows and exact-list or allowlist (StringNotEquals) denies accurate regardless of wire encoding. * s3api: preserve upload ACL denies and align policy checks Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> * s3api: retain upload owner grants and literal policy values Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> --------- Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
9d1c24d80d |
[Filer] Support append to inline small files (#11591)
* fix 11586 * Update filer_server_handlers_write_autochunk.go * filer: fix inline append races, empty files, and stale ETags Serialize the append read-modify-write on the entry lock so concurrent appends merge instead of losing content, keep small appends to empty files inline, tolerate legacy entries whose metadata size differs from their content, and set the entry digest so appended inline files keep a real ETag. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
76ddde6a6d | docs: regenerate star history chart | ||
|
|
0d93dec145 | helm: mount TLS certificates in bucket hook (#11589) | ||
|
|
d9b69a7f76 |
s3api: fix PutObjectAcl permission scoping and owner grants (#11587)
PutObjectAcl had four authorization and ownership bugs: - The handler embedded the resource path into the action (WriteAcp:bucket/object), and authRequest/CanDo then scoped it to the request's bucket/object again. A bucket-wide WriteAcp:bucket grant could never match, so legitimate owners got 403. - After authRequest succeeded via an IAM or bucket policy, a leftover identity.CanDo gate re-checked only the legacy Actions list, denying identities authorized purely by policies. - For canned and default ACLs, ExtractAcl generated the FULL_CONTROL grant for the requesting account instead of the object owner. An admin setting private/public-read on another account's object left the owner metadata intact but reassigned full control to the admin. - Objects without stored owner metadata (e.g. written via the filer outside S3) fell back to treating the requester as the owner, so any user with a WriteAcp grant could take them over. Non-admins are now denied; admins keep the takeover fallback. Grantee validation now also accepts the object's stored owner even when that account has been removed from the registry, so canned/XML ACLs for retired owners keep working. |
||
|
|
2b5fdc639f |
filer: stop isSameChunks from sorting caller-owned chunk slices (#11584)
* filer: stop isSameChunks from sorting caller-owned chunk slices slices.SortFunc reorders the input in place. filer.remote.sync calls IsSameData on a metadata event's NewEntry inside isMetadataOnlyUpdate and later stamps the filer entry under an IF_ENTRY_EQUAL precondition carrying that same entry. The ETag-sorted chunk list never matches the stored entry, so every stamp of a multi-chunk object fails, synced_mtime_ns stays zero, and dirty objects are re-uploaded forever. Sort clones of the slices instead. * filer: test IsSameData leaves input chunk order unchanged Guards the clone-then-sort fix: a regression back to in-place sorting would reorder caller-owned chunk slices and reintroduce the remote-sync IF_ENTRY_EQUAL mismatch. |
||
|
|
d7a02567e3 | docs: regenerate star history chart | ||
|
|
39bc9cd0ef |
s3api: copy the trailer checksum before reading the next trailer line (#11583)
* s3api: copy the trailer checksum before reading the next trailer line
parseChunkChecksum kept the checksum value as a sub-slice of the line
returned by bufio.Reader.ReadSlice, which is only valid until the next
read. When the trailer lines arrive in separate TCP segments, reading
x-amz-trailer-signature refills the buffer and overwrites the saved
value, so a correct upload fails with InvalidDigest ("The Content-Md5
you specified is not valid").
The AWS SDK for Java v2 (>= 2.30) on a Linux JDK sends the trailer that
way; about half of its signed streaming uploads failed.
Fixes #11582
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* s3api: reuse crc32 writer and trim comments in trailer split test
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
|
||
|
|
f4ef37e752 |
filer.sync: resubscribe the metadata stream when a failure pins the offset (#11581)
* filer sink: keep the gRPC status inside wrapped errors
%v stringifies the status, so a peer teardown reported as Canceled ("the
client connection is closing") reached IsTransientError as plain text and
matched nothing: the sync job failed on the first attempt and pinned the
offset. %w keeps the status reachable, so the retry runs on a fresh
connection once the target is back.
* pb: let a consumer drop the metadata stream to force a resubscribe
A MetadataProcessor job that exhausts its retries pins the processed
watermark so the event replays on the next subscribe — but nothing on the
source stream notices a target-side failure, so the replay waited for an
unrelated reconnect or a restart. The new Resubscribe channel cancels the
stream's context; the Recv loop answers it with ErrResubscribe so the
caller's retry loop resubscribes from GetResumeTsNs and replays the pinned
events in order.
* pb: stop the event retry loop once the stream context is done
RetryUntil ignores context, so a subscriber parked on a failing offset
write would keep retrying past a resubscribe signal until the sink came
back. Stop retrying when the stream is being dropped so the resubscribe
takes effect promptly.
* filer.sync: signal resubscribe when a job failure pins the offset
A job that exhausts its in-job retries leaves the event pinned behind oldestFailedTsNs, replayable only on a reconnect. Closing resubscribeCh on the first recorded failure lets the metadata follower drop the stream so the reconnect replays the pinned events instead of waiting for a process restart (#11572).
* filer.sync: wire the resubscribe signal into the follow options
filer.sync, filer.remote.sync, and the remote gateway bucket sync all run their subscription inside an outer retry loop, so ErrResubscribe resurfaces as a resubscribe from the persisted watermark.
* filer.sync: wait for in-flight jobs before signaling resubscribe
* remote sync: never resume past the saved offset when -timeAgo is set
* filer.sync: drop events that arrive after the drain signals resubscribe
* pb: interrupt the event retry backoff when the stream context ends
* filer.sync: stop admitting once a failure pins, and count jobs per timestamp
A pinned watermark only released once the processor went fully quiet, so a busy stream could starve the resubscribe — the failed event would wait for an unrelated reconnect anyway, the wait this mechanism exists to remove. The processor now latches stopped when a job fails: admission drops new events (they replay from the pinned watermark after the reconnect), a broadcast releases blocked waiters, and the resubscribe signals as soon as the jobs already in flight drain. A redelivery of an event still in the failure ledger may still run so its success shrinks the replay, but nothing starts once the signal has fired, or it would race the replay it asked for.
Dropped events no longer inflate the received counters — an event counts only once admitted, and the replay's own admission counts it.
While here: activeJobs keyed by TsNs collapsed events sharing a timestamp, so one completion could empty the map while a same-ts sibling was still running — letting the drain gate and the watermark outrun it. Jobs are now counted per timestamp, and the drain and lazy heap cleanup go through the counts.
|
||
|
|
10b0f2b8ad |
volume server: refuse the rest of a grouped run after a durable index failure (#11576)
* volume server: refuse the rest of a grouped run after a durable index failure A durable write whose needle-map put fails stops the volume taking writes (#10825): sent on its own, the next write then fails read only before it appends. The grouped run from #11543 appends and syncs every entry before publishing any, then kept publishing the entries after the failed one and acked them once the shared .idx sync went through. When the failed put tore its .idx row, the rows appended after it land off alignment, so the next load parses them as garbage and the acked writes are gone. Once a durable entry fails to publish, refuse every later entry of the run with ReadOnly, as the per-needle path does. The entries before it stay acked; their rows go down with the run's one .idx sync. The refused records stay on the .dat unindexed, as the failed one does on its own. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: refuse a grouped entry staged as a cookie mismatch too After a durable entry in a grouped run fails to index, the entries after it are refused as they would be on their own. On its own an entry meets check_writable before its cookie check, so one staged as a cookie mismatch now gets the refusal too, instead of keeping its staging error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: trim a torn .idx row back so the next stays aligned A failed write_index_entry can leave half a row in the .idx. With the writer appending at the tail, every row written after it lands off alignment and the next load parses them as garbage, so a write acked behind a torn row does not come back. Trim the file back to idx_file_offset on a failed append, in both needle maps, and cover it with a test that writes past a torn row and reloads. * volume server: refuse queued Go writes once a durable index update fails processBatch kept writing after a failed nm.Put, and the single-write path checked IsReadOnly only outside the volume lock. A durable write whose index update fails now marks the volume noWriteOrDelete, and each queued request is checked before it appends, so the ones after a failed durable entry are refused the way a lone write is. Deletes get the same noWriteOrDelete refusal a lone delete gets. * volume server: refuse appends while a torn .idx row cannot be trimmed When trimming back a half-written .idx row itself fails, the next append would land after the torn bytes and every later row would parse off alignment on load. Latch the map as torn and refuse appends until the trim succeeds, on both CompactNeedleMap and RedbNeedleMap; the same latch covers an orphan row that could not be trimmed after a failed redb commit. The .idx writer is now opened with write+append access so truncate_to (set_len) works on Windows, where an append-only handle cannot trim. * volume server: write .idx rows at idx_file_offset, not via append mode Rust's OpenOptions on Windows strips FILE_WRITE_DATA whenever append is set so the handle stays strictly append-only, which makes set_len fail - the torn-row trim could never succeed there. Open the .idx writer with plain write access and seek to idx_file_offset before each row, the same positioned-write model the Go server uses. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
eafe79ebff |
filer: skip UpdateEntry when inline content is unchanged (#11580)
* filer: skip UpdateEntry when inline content is unchanged SaveInsideFiler rewrites config files (IAM identities, filer.conf, remote mappings, policies) unconditionally. Each no-op UpdateEntry is a metadata event the local meta log persists to /topics/.system/log, which appends a chunk to a volume. A client that rewrites identical config on a timer, e.g. the seaweedfs-operator 5-minute resync calling UpdateUser with unchanged actions, keeps .dat/.idx files growing on an otherwise idle cluster and prevents HDD spindown (seaweedfs/seaweedfs#11571). Skip the UpdateEntry when the stored inline content is byte-identical, so unchanged writes produce no metadata event and no volume writes. * filer: test that identical SaveInsideFiler writes skip UpdateEntry * filer: require stamped Md5 before skipping identical writes An entry holding identical content but no Md5 (written before hashing, or by a tool that cleared it) would never get the stamp that IF_ETAG_MATCH conditional writes key off. Skip only when both the stored content and its Md5 match, so one write still lands to repair the stamp. |
||
|
|
562afa8ec9 |
filer: resume metadata subscriber from processed watermark on reconnect (#11574)
* filer: resume metadata subscriber from processed watermark on reconnect
* filer: take the reconnect position from GetResumeTsNs verbatim
The callback is the subscriber's durable resume point; falling back to
StartTsNs when it returns zero can resume from a cursor the log-chunk
reader advanced past still-pending work.
* filer: advance the stream cursor once a retried event recovers
RetryForeverOnError resolves the failure inside handleErr, so returning
without moving StartTsNs replays work the event already did when the
stream reconnects before the next one arrives.
* filer: let filtered-progress markers move the processed watermark
A marker means the source examined everything up to its timestamp and
skipped what did not match the subscription. With a resume callback the
marker now reaches the consumer, and AddSyncJob advances the watermark
to it once every earlier job finished and no failure pins the offset.
Idle filtered stretches no longer rescan on every reconnect, while the
guards keep the watermark behind pending or failed work.
* filer: unpin the watermark once a failed event completes
oldestFailedTsNs was only ever set, so a failure that a replay later
fixed still held the resume offset, and every reconnect re-read the
same backlog. Track outstanding failures in a set and recompute the
pin when the failed event's job finally succeeds.
* filer.remote.gateway: resume bucket sync from the processed watermark
The bucket-sync subscriber runs the same MetadataProcessor queue as
filer.remote.sync; give it the same GetResumeTsNs callback so a
reconnect resumes from durably processed work, not the last seen event.
* util: treat a peer-sent gRPC Canceled as transient
A peer tearing down its end of the transport reports codes.Canceled
("the client connection is closing"), which IsTransientError used to
reject: the sync job then failed on the first try and held the offset
until a restart. Caller's own cancels are still excluded up front by
errors.Is(err, context.Canceled), so only teardown-style statuses take
the new branch.
* fix: preserve filtered progress and distinguish caller cancellation
* filer: bound the failed-event ledger past a persistent outage
A destination rejecting every event grew failedTs by one entry per source
event for the life of the processor. Past maxFailedSyncEvents the set now
collapses to a sticky pin at the smallest failure seen, so the watermark
still replays from the oldest failure while memory stays bounded; a
restart re-derives the exact set.
Also keep a resume-callback consumer's chunk-ref replay filter at the
subscribe-time position instead of option.StartTsNs, so a resubscribe does
not filter out events whose async processing is still pending.
* filer: key the failed-event ledger by event, not just timestamp
A success for one event cleared the pin recorded for a different event
that shared its TsNs, letting the watermark pass an unresolved failure.
The ledger now keys on the event's path identity, so recovery unblocks
only the event that actually failed.
---------
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
|
||
|
|
8c1ebbee32 |
volume server: split get_or_head_handler_inner into phases (#11489)
* volume server: read GET/HEAD needles off the store lock, and only once The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stream needle chunks without the store lock StreamingBody::poll_frame took store.read() and find_volume for every chunk to compare the volume's compaction revision, dup'd the source handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the stream also holds a data-file read lease for its whole life, while a writer waits for that lease under store.write(): the next chunk's store.read() then waits for the writer and the writer for the stream. The per-chunk re-lookup was also wrong. The stream reads a handle opened at plan time, which pins the .dat inode the offset was resolved against; a vacuum commit renames a new file over .dat and leaves that inode untouched. The re-looked-up offset belongs to the new file but was read from the old inode, so a stream whose needle a vacuum moved ended in a checksum error. The pinned offset stays valid, so the check, and with it every store access, is dropped, along with the now unused re_lookup_needle_data_offset and the revision fields of the read plan. The source is shared as an Arc instead of dup'd per chunk, and the chunk buffer is a BytesMut that the blocking read hands back with its result, so its allocation is reclaimed once the previous frame has been written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: split get_or_head_handler_inner into phases get_or_head_handler_inner was a ~650-line function. Its middle resolved the needle and set five mutable flags (stream_info, can_stream, can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that three if-let reply paths then re-tested, each re-checking stream_info. It is now a 126-line orchestrator over named phases: reject_read_jwt, proxy_missing_volume, wait_for_download_slot, parse_read_request, read_ec_needle / read_volume_needle, etag_and_last_modified, not_modified_response, read_response_headers, and the reply phases stream_response, head_from_meta_response, range_from_source_response, buffered_payload and buffered_response. The read phases return a ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource, Buffered) carries the NeedleStreamInfo only on the variants that use it, so the reply is one match instead of three flag checks. Pure refactor: every status code, header and header order, error text, metric increment, lock and data-file lease scope, spawn_blocking boundary and side-effect order is unchanged. Phases that can end the request return ControlFlow<Response, T>. A Range header that is not visible ASCII still falls through to the buffered path, as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a needle stream once its volume becomes unavailable Taking the store lock out of StreamingBody also dropped its per-chunk unavailable_error() check. With -hasSlowRead a writer can take the data-file lease between chunks, fail its fsync and its truncate, and mark the volume unavailable; the stream then kept serving the rest of the needle from its pinned handle. The volume's io_unavailable reason is now an Arc-shared leaf mutex that the read plan hands to the stream. Each chunk checks it under its data-file lease, where the writer marks it, and fails with the same "volume is unavailable: <reason>" error the old check returned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: mirror Go order in the buffered read path - check HEAD before Range in buffered_response (writeResponseContent order); an EC-volume HEAD with a Range header answered 206, Go answers 200 - treat the proxied flag as an exact query pair like Go's parsed lookup, not a substring - name the phases after their Go counterparts: check_download_limit and read_ec_shard_needle; reuse has_replication() - drop comments that restate the code or cite Go line numbers --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
c9ade3f9fd |
helm: compare PVC sizes numerically in the volume resize hook (#11575)
include always returns a string, so gt compared the rendered quantities (e.g. "1.2884901888e+11" vs "6.442450944e+10") lexically. Growing a volume from 60Gi to 100Gi/120Gi or 500Gi to 1Ti emitted no kubectl patch: the StatefulSet was recreated with the new volumeClaimTemplate but the PVC kept its old size. Shrinks such as 120Gi -> 60Gi emitted a patch instead. Pipe both values through float64 before comparing. |
||
|
|
6c07a5fdd0 |
s3: keep small ranged GETs on range reads, no whole-chunk downloads (#11577)
* filer: keep a ranged read in random mode through its contiguous tail A far ReadAt on a fresh ReaderPattern left the sequential counter at -1, so the next buffer of the same ranged request landed on the frontier and flipped the verdict straight back to sequential — readChunkSliceAt then paid a whole-chunk fetch for the remainder of the range. Drop the counter to -ModeChangeLimit when random mode is entered so the verdict needs sustained sequential evidence to undo, matching the hysteresis an established sequential stream already gets. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: pin small ranged GETs to range reads A ranged GET whose first read lands within SeqTolerance of offset 0 is judged sequential immediately, and even a far-starting range could flip back mid-request; either way readChunkSliceAt downloads each covered chunk in full, multiplying disk reads for small ranged reads (measured ~7x). Pin random mode for ranged requests no larger than SeqTolerance so all of the request's buffer reads stay range fetches. Larger ranges keep the dynamic pattern, where whole-chunk fetches amortize. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: fetch only the part of a chunk the view covers Replaces the PinRandomMode size heuristic with a per-chunk coverage rule. ViewFromVisibleIntervals already clips chunk views to the request window, so a view that is not IsFullChunk() is one the request only partially needs; fetch it as a range regardless of the detected read pattern. This closes the holes a request-size pin left open: ranges larger than SeqTolerance no longer revert to whole-chunk downloads once their buffers look sequential, and ranges that fully cover a chunk keep the shared whole-chunk path instead of fetching 256KiB slices piecemeal. Prefetch (MaybeCache) skips clipped views so it cannot amplify a range read either. PinRandomMode is dropped: no caller needs it once coverage drives the fetch choice. Range fetches route through fetchChunkDataFn so tests observe them the same way as whole-chunk downloads. * filer: keep ciphered chunks on the whole-chunk path A range fetch cannot save bytes for a ciphered chunk: readEncryptedUrl always downloads and decrypts the whole blob before slicing. Sending partial views of ciphered chunks through fetchChunkRange would repeat the full download per buffer, so they keep the shared whole-chunk path where one download serves every buffer. Prefetch stays enabled for them for the same reason. * filer: keep compressed chunks on the whole-chunk path Like ciphered chunks, a range request on a compressed chunk makes the volume server read and decompress the whole needle, so range-per-buffer would repeat the full backend read for each 256KiB window. Route them through the shared whole-chunk path via ChunkView.CanRangeFetch. * filer: fall back to range fetch when a chunk exceeds the reader budget A ciphered or compressed chunk larger than readerCacheSizeMB can never be read through the whole-chunk path — the budget rejects the buffer — so its partial views must still range-fetch or the GET fails outright. --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
07da302da0 |
volume server: ec.decode verifies, cleans up and compacts like Go, off the runtime (#11547)
* volume server: ec.decode reads the .ecx from the index dir it was copied to VolumeEcShardsCopy writes the .ecx/.ecj into the receiver's -dir.idx, so with a split data/index dir the decode target has no .ecx beside its shards. VolumeEcShardsToVolume sized the .dat from the right .ecx but built the .idx from the data dir, failing with NotFound after the .dat was already published. It now reads .ecx/.ecj from where the EC volume opened them and writes the .idx beside the .dat, where Go leaves it. The live-entry check and the .dat size also ignored deletions recorded only in the .ecj, which Go folds into the .ecx (RebuildEcxFile) first: a fully deleted volume was decoded instead of reported as having no live entries, and deleted tail needles were copied into the .dat. Both now treat journaled ids as deleted, without rewriting the sealed .ecx. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode keeps the decoded volume writable and reads every .ecj The rebuilt .idx copied a journaled tail needle's .ecx row verbatim after the .dat was cut short before it, so the mount saw a row past EOF and marked the decoded volume read-only. Rows of deleted needles the .dat no longer holds are now dropped, and each journaled needle still in the .dat gets one tombstone instead of one per journal entry. VolumeEcShardsCopy appends journals collected from other holders into the idx dir, but the decode read only the .ecj beside the .ecx, which sits in the data dir when this server generated the shards. It now reads both, once, in bounded chunks via the loader EcVolume uses. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: test ec.decode drops a sealed .ecx tail tombstone Covers the other half of the rule added in the previous commit: a tail needle tombstoned in the .ecx itself (Go's RebuildEcxFile) is cut from the .dat, and its row must not reach the rebuilt .idx either. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode runs its file I/O off the async runtime VolumeEcShardsToVolume released the store lock before decoding, but read the .ecx/.ecj, rebuilt the .dat and wrote the .idx inside the async handler, parking a runtime worker for the length of a volume-sized copy. The decode now runs in spawn_blocking on inputs snapshotted under the store lock. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode checks the rebuilt .dat is complete Go stats the decoded .dat before writing the .idx (VerifyDecodedDatFile) and fails the decode when it is shorter than the extent the EC index references, since the caller deletes the shards once the call returns. The Rust handler returned success without that check. The rebuild already fails on a short shard read, so this guards the published file itself. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode drops the decoded volume's bitrot sidecars Go removes <base>.ecsum and <base>.ecsum.v<N> beside the .dat and beside the .ecx once the .idx is written, so a stale checksum sidecar cannot pass for the protection of a later re-encode. The Rust handler left them in place. Removal is best effort, as in Go. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode compacts the decoded volume Go ends VolumeEcShardsToVolume with an offline CompactVolumeFiles, so the decoded volume holds only live needles. The Rust decode left every needle deleted through the .ecj in the .dat, tombstoned in the .idx, until a later vacuum reclaimed it. Store::compact_volume_files loads the unmounted volume, checks free space the way the vacuum does (the estimate now lives in one helper), and runs the vacuum's compact-by-index and commit. As in Go a failed compaction is logged and the decode still succeeds, so the uncompacted .idx rules stay: the tests that pin them now make the compaction fail. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: ec.decode keeps deletes journaled while the .dat is written The decode read the .ecj journals once, before rebuilding the .dat, so a delete that reached the EC volume during the rebuild was left out of the new .idx and the needle came back live. Each journal's read length is now kept, and the bytes appended since are read just before the .idx is written, after waiting out any journal append in flight (appends hold the store write lock), so every delete acknowledged by then is in the .idx. A delete after that point is still lost, as in Go. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Guard overlapping ec decode requests; serialize journal catch-up volume_ec_shards_to_volume runs its decode in spawn_blocking, so a dropped request leaves the job running and a retry would race it on the temporary and final volume files. Claim the vid in a per-server in-flight set until the blocking job finishes, and return Unavailable to an overlapping request. The Go handler has the same exposure and gets the same guard. Journal appends hold the store write lock through their sync-or-truncate, so holding a read lock across the catch-up read guarantees every record it sees is committed: a rolled-back delete can no longer leave a tombstone in the decoded index. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Reconcile the swap when offline compaction commit fails A CommitCompact that fails after the .cpc marker may have renamed .dat but not .idx. cleanup_compact refuses while the marker exists, so the mismatched pair survived until a restart reconciled it — and the decode caller treats the failure as non-fatal. Run reconcileCompactState on commit failure so a decided swap rolls forward and orphan temps are removed before the volume can mount. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Release the decode claim on panic * volume: add ec_decodes_in_flight to the integration-test state literal * volume server: hold the decode tail's lock through compaction The catch_up read released before the rebuilt .idx was written and the volume compacted, so a delete synced to .ecj in that window was durably journaled yet absent from the published index — resurrecting the needle. Rust now holds the store read lock from catch_up through compact, and Go mirrors it by holding the volume's journal lock from the journal- consuming index write through CompactVolumeFiles. * volume server: serialize ec decode's tail per volume, not per store Review follow-ups on the decode path: - Rust: holding the store read lock from journal catch-up through the offline compaction stalled every writer on unrelated volumes for the whole rewrite. The new ec_decode_tail set marks the vid only while its .idx is published and .cpd/.cpx swapped; the two local .ecj append paths (VolumeEcBlobDelete, the distributed delete's local journal) wait on a Notify for that span — Go's per-volume ecjFileAccessLock semantics without the global stall. VolumeMount and the staged-adopt path are also held off while a decode claim is in flight so neither can race the swap. - Rust: the initial journal read ran unlocked, so bytes a rolled-back append later truncated could be folded in as phantom tombstones. The first pass stays unlocked (a slow journal must not stall the store) and a rescan under the quiescing read lock re-reads only committed content; catch_up now rebuilds the id set when a regular journal shrank. - Go: the decode resolved the compaction DiskLocation through FindEcVolume while holding the journal lock, inverting DestroyEcVolume's map->journal order into a deadlock. The lookup now happens first, and DestroyEcVolume/deleteEcVolumeById/DiskLocation.Close destroy outside the map lock. - Go: RebuildEcxFile unlinks .ecj while the volume's ecjFile handle stays open, so later deletes could commit to a detached inode. Both call sites now fold under the journal lock and ReopenDeletionJournal repoints the handle at the live path, working on the volume's resolved .ecx dir (EcIndexBaseFileName) rather than the configured index dir. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: fence EC remounts behind the destroy tombstone DestroyEcVolume, deleteEcVolumeById, and the collection-delete sweep now remove the EcVolume from ecVolumes before destroying it off-lock, so a concurrent remount could re-open shard files that the in-flight destroy then unlinks — registering a detached fd. Each destroy records a per-vid tombstone channel in a new ecVolumesDestroying map before dropping the map entry and closes it when Destroy returns. The tombstone intentionally survives as the vid's destroy generation: loadEcShardWithIdxDir compares it before and after opening the shard, so a destroy that both started and finished inside the open window is still detected. A mismatch drops the just-opened shard (releasing its fd and mount gauge) and retries after the destroy completes; a successful mount clears the stale tombstone. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: rescan the .ecj under the store lock only after a rollback The decode's second journal pass ran a full rescan under the store read lock on every decode, stalling unrelated writers for the length of the scan. Bump a process-wide epoch whenever a failed append truncates its uncommitted tail; an unchanged epoch between the unlocked read and the quiesced pass proves every id folded in was committed, so catch_up() suffices. catch_up() also treats a journal that was read but has since disappeared as shrunk to zero, so its earlier ids cannot linger. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: check the decode tail under the store write lock on delete A blob delete waited for the publishing tail before taking the store write lock, so a decode that claimed the tail while the delete was parked behind the decoder's read lock could still see the journal append land after the rebuilt .idx — an acknowledged delete the mount would miss. Test tail membership under the write lock instead, retrying after the wait; journal_delete_local reports WouldBlock for the same recheck on the distributed path. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: claim the vid for mount and staged adoption, per volume VolumeMount and the staged .copying adoption held the ec_decodes_in_flight set lock through slow file renames and mounts, stalling every unrelated volume's decode, mount, and adoption. Take the per-volume claim instead — the same exclusion against a racing decode for this vid, released when the call returns. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: fail the decode when a compaction commit marker survives CompactVolumeFiles' caller logged a compaction error and went on to delete the EC shards. When the commit marker (.cpc) is still on disk the .dat/.idx swap was decided but could not be reconciled, so the mounted pair may be mismatched — report the failure instead so the shards are kept and the caller can retry. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: gate the parked-delete test on the held write lock The releaser thread and the spawned delete raced for the store write lock; on a slow runner the delete could acquire it first and commit before the tail was ever claimed, failing !delete.is_finished() on the Windows unit-test job. Spawn the delete only after the thread reports the lock held. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
68df7511f6 |
filer.remote.sync: do not pin the sync offset on completed work (#11569)
* filer.remote.sync: do not pin the sync offset on completed work * filer.remote.sync: a superseded rename uploads the current entry; typed NotFound for a stamp on a deleted entry * filer.remote.sync: a superseded rename keeps the old key when it is the only copy and uploads once * filer.remote.sync: a rename whose content is now remote-only fails the event instead of completing it * filer.remote.sync: a remote-only rename copies the old object to the destination before deleting it * filer.remote.sync: the remote-only rename path follows the filer's current entry and verifies the destination object * filer.remote.sync: an event that described an entry without data is superseded once the filer wrote to it * filer.remote.sync: a superseded rename does only the work left to do uploadCurrentEntry met a remote-only current entry with a fixed error, but a sync plus remote.uncache in the meantime leaves the destination holding the stamped object; that state is complete, not lost. The remote-only case now finishes through completeRemoteOnlyRename, which verifies the destination against the entry stamp and fails only when neither key holds the content. A current entry whose stamp covers its content was already uploaded by the superseding event; skip it instead of writing the same bytes again. * filer.remote.sync: an inherited stamp does not prove the content synced The stamp-coverage skip in uploadCurrentEntry read LastLocalSyncTsNs as proof the current content was uploaded, but a rename carries the source entry's stamp to the destination: a rewrite hidden by that stamp (the case the fallback upload exists for) carries a LastLocalSyncTsNs at or after its mtime and would have been skipped. Drop the check; the remote-only path verifies content at the destination itself through describes. --------- Co-authored-by: James Sas <james@medable.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
793ce06b10 |
s3api: allow unsigned SSE-C customer key headers on presigned requests (#11578)
* s3api: allow unsigned SSE-C customer key headers on presigned requests AWS requires only x-amz-server-side-encryption-customer-algorithm to be signed on presigned URLs; the key and key-MD5 headers are supplied at request time. Since #9121 rejected any x-amz-* header outside SignedHeaders, SDK-generated presigned SSE-C requests (e.g. .NET GetPreSignedUrlRequest) fail with SignatureDoesNotMatch. Exempt the customer key and copy-source key headers for presigned requests only. * s3api: test presigned SSE-C requests carrying unsigned key headers |
||
|
|
0ca484c354 |
vacuum: bound master vacuum RPCs with phase deadlines (#11579)
* vacuum: bound the commit RPC with a phase deadline VacuumVolumeCommit ran on context.Background(), so a volume server that keeps the call pending would hold the topology-wide vacuum guard forever and every later sweep would be skipped. Give the call a deadline scaled like the existing phase waits (one minute per GB of the volume size limit) so a stalled commit ends as an error instead of blocking the sweep; the timeout is a var so tests can shrink it. * vacuum: bound the replica status probe with a phase deadline The VolumeStatus call on replicas that were not compacted also ran on context.Background(), so a stalled replica could pin the sweep the same way a stalled commit can. Give it the same per-phase deadline. * vacuum: bound the cleanup RPC with a phase deadline VacuumVolumeCleanup also ran on context.Background(); a stalled server would keep the sweep worker and the shared vacuum guard pending forever. Give it the same per-phase deadline. * vacuum: let the check and compact phase waits cancel their RPCs The coordinator wait timers fired while the check and compact calls still ran on context.Background(), so the sweep gave up but the RPC goroutine stayed until the server answered, and a compact stream kept writing on the server. Share one deadline context between the wait and the calls so an expired wait actually cancels them. * vacuum: test that a stalled volume server releases the vacuum guard A fake volume server keeps one vacuum-phase RPC pending until the client context is cancelled. Before the phase deadlines, Vacuum never returned and vacuumLockCounter stayed held; now each phase cancels on its deadline and the guard is free for the next request. * volume: stop compaction at the next needle when the client cancels The progress callback only noticed a gone client when a 128 MiB report failed to send, so an aborted VacuumVolumeCompact kept copying for up to a whole interval while the master had already moved on to cleanup. Check the stream context on every needle, the same early return the Rust volume server does with tx.is_closed(). * vacuum: assert the stalled phase RPC is cancelled, not just bypassed The check and compact coordinator waits already returned on timeout before the deadlines existed, so a regression that put the calls back on context.Background() would pass unnoticed. Wait for the fake server to report that the phase RPC context ended. * vacuum: give the stalled-RPC test room to reach the handler The 50ms phase budget starts before goroutine scheduling and the gRPC dial, so a busy test host could expire it before the fake server saw the call. Raise the override to 250ms; the test still finishes in about a second. * vacuum: describe the phase deadline as scaled, not per-GB The formula keeps the exact expression the check and compact waits already used (floor plus one at 1 GiB granularity); it is a backstop, not a per-GB SLO. |
||
|
|
3e679e925e |
filer: keep generated inodes inside the positive signed 64-bit range (#11567)
AsInode derives inodes from HashStringToLong, which is uniform over int64, so roughly half of the derived values land above math.MaxInt64 once they are converted to uint64. The Elasticsearch store indexes Entry.Attr.Inode as a signed long, so those values are rejected with HTTP 400 and the metadata entry is never written, which the filer then retries forever. Fold the sign bit off in one place, util.NormalizeInode, and route both derivation sites through it: FullPath.AsInode (path plus creation time) and the hard-link branch in ensureEntryInode (HardLinkId hash). Masking keeps the other 63 hash bits, so distinct paths still get distinct inodes, and it applies identically to the FUSE mount, which derives the same value. Co-authored-by: Yi-111-a <34116709+0-xiaosu@users.noreply.github.com> |
||
|
|
52fb9f93ff |
s3: track filer joins and leaves pushed by the master (#11563)
* fix(s3): track filer joins and leaves pushed by the master The S3 FilerClient replaced its -filer seed with a master snapshot of filer IPs at boot and refreshed it only every 5 minutes. A rolling restart replaces every filer well inside that window, leaving S3 servers with only dead addresses and failing every write until the next poll. Apply the master's ClusterNodeUpdate pushes to the filer list as they arrive, keeping the poll as a backstop. The last filer is never removed, and a poll snapshot requested before a push was applied is discarded rather than overwriting newer membership. * Defer last-filer leaves; bump the generation only on real changes * fix(s3): cancel deferred filer leaves on rejoin and on discovery A deferred last-filer leave outlived the filer it was recorded for: a rejoin at the same address looked like a duplicate add, and a discovery snapshot left the entry behind. The next join then removed a live filer until the following poll. A join now cancels any deferred leave for its address, and an applied snapshot clears them, since it is the master's current membership. * Bump the push generation when a rejoin cancels a deferred leave --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
0ff7794c54 |
volume: compact an oversized .ecj at mount, safely (Rust + Go) (#11555)
* volume: compact an oversized .ecj at mount, safely (Rust + Go) Restore the mount-time compaction dropped from #11408, Rust + Go parity. A journal already bloated by repeated shard copies is folded down to the id set it encodes. - Trigger after load when file_records > max(threshold, 4x distinct), with a 1 MiB floor so small journals are never rewritten. The set is written to .ecj.compact.tmp + fsync, the handle dropped, renamed, the directory fsynced and the append handle reopened. A failure before the rename keeps the original journal and handle; a failure after it fails the mount. - Go never compacts after a failed journal load; the set would be partial and the rewrite would drop the unread records. - A per-path registry (ecj_registry.rs / ecj_registry.go) counts EcVolume holders and out-of-band writers of each .ecj. Compaction runs only when this volume is the sole holder and no copy is writing; holders and writers wait while one runs. This covers shared -dir.idx journals and cross-disk reconcile, where another EcVolume may hold the same journal. - VolumeEcShardsCopy and EC index recovery register as writers around their .ecj append and partial-file cleanup. - Under the reservation, re-check that the file on disk is still the inode and size that was loaded. - Publish errors are classified where they happen; a failed rename plus a failed restore reports both errors. - Compaction runs after the .vif / bitrot checks, so a refused mount leaves the journal untouched. - The tmp is opened like other volume files, removed at mount if a crash left it, and listed in every EC index cleanup path. Failure paths are tested through the real mount via injectable fs steps (open_with / newEcVolumeWith), plus sibling holders, active copies, changed-after-load, stale tmp cleanup, refused mounts and the Go load-error guard. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: fail the mount when the compacted .ecj's directory cannot be synced The Rust mount synced the journal's directory after renaming the compacted file over it through the crate's best-effort fsync_dir, which returns Ok when the directory cannot be opened. A rename needs only write and search permission, so on a directory without read permission the replacement was published, never synced, and the mount went on taking deletes against it. Sync through a helper that propagates the open error, as Go's util.FsyncDir already does, so that case fails the mount like any other post-rename sync failure. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: test the no-compaction-after-failed-load rule through the Go mount The test for it handed compactEcjAfterLoad an artificial error on a volume that had loaded cleanly, so it would not notice NewEcVolume dropping the real load error on the way to compaction. Make the journal read one of the injectable ecjFsOps steps and fail it inside the real mount, after the first chunk, on a journal whose last entry is an id the first chunk does not hold. The mount must leave the file byte for byte as it was; a clean remount then compacts and keeps that id. The Rust mount fails outright on a load error, so it has no equivalent path. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: register ReceiveFile's .ecj writes with the journal registry ReceiveFile refuses a mounted EC volume only once, when the info message arrives, then creates the .ecj and streams chunks into it. A volume that mounted on that journal mid-stream could find a bloated prefix, pass the inode-and-size re-check and rename a compacted file over it; the rest of the stream then went to the unlinked inode and was lost. Register the path as a writer before the file is created, in both the Go and Rust handlers, and hold it until the file is closed and any partial copy removed, as the shard-copy and index-recovery appends already do. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: skip .ecj compaction when a writer ran since the journal was loaded Compaction checked only that no writer was active at the reservation, and that the file was still the loaded inode at the loaded size. A ReceiveFile truncates and refills the journal in place, so one that ran during the mount's load, or after it, and finished before the reservation could leave different ids at the same length; compaction then wrote the stale set over them. Give each path a write generation that every writer bumps as it starts. A holder records it, and whether a writer was active, when it registers, which is before it opens and loads the journal. It may compact only if no writer was active then and the generation has not moved. Same rule in Go and Rust; the journal read becomes an injectable step in Rust as it is in Go, so both test the in-place rewrite through the real mount. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: match the ReadOnly(VolumeId) variant in write_volume_needles #11543 matched VolumeError::ReadOnly as a unit variant in Store::write_volume_needles, and #11544 changed it to ReadOnly(VolumeId) in the same merge window. Each passed CI on its own, but master no longer compiles the Rust volume server. Carry the volume id through. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
94edd0a6d4 | docs: regenerate star history chart | ||
|
|
30069f3e45 |
iam: manage OIDC providers and roles over the filer IAM gRPC service (#11523)
* s3/iam: manage roles through the IAM API, with an opt-in persistent role store
Roles could only come from the IAM config file: the S3 server pinned the
role store to memory and the embedded IAM API had no role actions, so a
role could not be created, retrusted or revoked without editing the file
and restarting every gateway.
Role store
- Read the `roleStore` key (the IAMConfig field already existed). With an
IAM config file the default stays memory; with none it is the filer, as
for OIDC providers, so zero-config clusters keep runtime-created roles.
- Roles from the IAM config file never go into a persistent role store,
which outlives the file and may be shared by S3 servers with different
files. They are served from memory beneath the store, as OIDC providers
are: a stored role of the same name takes precedence, and deleting it
restores the file's. A config-file role cannot be changed or deleted
through the API (UnmodifiableEntity), and removing one from the file
removes it at the next start. An in-memory store holds them as records,
as before. They have no creation time, so CreateDate is omitted rather
than reporting when this server started. SetRoleStore installs a store
the same way, so a store set after startup keeps the config-file roles,
as SetOIDCProviderStore does for providers.
- Watch /etc/iam/roles and drop the cached role definitions on change. The
cached filer store otherwise serves a peer's stale role for up to its 5m
TTL, which keeps a revoked trust policy in force on the other gateways.
- Role stores wrap ErrRoleNotFound for a missing role; the filer store
used to report any failed lookup as "role not found". CreateRole proceeds
only on a confirmed absence, so an unreadable store cannot let it write
over an existing role.
IAM actions
- CreateRole, GetRole, ListRoles, DeleteRole, UpdateAssumeRolePolicy,
AttachRolePolicy, DetachRolePolicy, ListAttachedRolePolicies. The reads
are allowed in read-only mode.
- A role defined in the config file is reloaded from it at every start, so
changing or deleting it through the API is refused (UnmodifiableEntity)
rather than silently reverted.
- DeleteRole with policies attached is refused (DeleteConflict), as on AWS.
- Role names follow AWS's rules ([\w+=,.@-]{1,64}); a role is stored as
<name>.json in the filer, so this also keeps a name from leaving the role
store's directory. At most 10 managed policies per role (AWS's default
quota; MaxManagedPoliciesPerUser is 10 too), LimitExceeded beyond.
- DeletePolicy is refused (DeleteConflict) while a role attaches the
policy, as it already is for users and groups: roles attach policies by
name, so a policy created later under the deleted one's name would
otherwise take effect on the role.
- Role paths other than "/" and role tags are not stored, so they are
refused rather than dropped.
Role IDs and sessions
- Roles get a unique RoleId when first stored (random, AWS AROA form),
kept across updates; a config-file role gets a stable ID derived from its
name, since it is created again at every start.
- Sessions issued through AssumeRoleWithWebIdentity, AssumeRoleWithCredentials
and AssumeRole carry the role's ID (claim "rid"), and a request under a role
whose current ID differs is denied. Resolving a session's policies by role
name let a session outlive its role: once a role was deleted, a role later
created under the same name — with a different trust policy and different
policies — revived every unexpired session of the old one with the new
role's permissions. Sessions issued before this change carry no ID and are
unaffected until they expire.
Integration test (test/s3/iam, run with `make start-services`):
TestWebIdentityWithProviderAndRoleManagedThroughIAMAPI configures an OIDC
provider, a managed policy and a role entirely through the IAM API against a
JWKS served by the test, then checks the trusted subject gets credentials
scoped to the attached policy; another subject, a token signed by another
key, an unsigned token and a token for another audience are refused; and UpdateAssumeRolePolicy moves the
trust at once.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* iam: manage OIDC providers and roles over the filer IAM gRPC service
The filer's SeaweedIdentityAccessManagement service covers users, access
keys, policies and service accounts, but not the OIDC providers and roles
that STS web-identity federation needs. A controller that already manages
IAM over this service (seaweedfs-operator's S3OIDCProvider) has no
transport for them; its swadmin client returns ErrOIDCNotWired and names
this as the recommended fix.
- PutOIDCProvider / GetOIDCProvider / DeleteOIDCProvider / ListOIDCProviders
and PutRole / GetRole / DeleteRole / ListRoles.
- They write the filer-backed stores at their default paths, which S3
servers read when configured with a filer-typed "oidcProviderStore" and
"roleStore"; the S3 servers' /etc/iam subscription applies changes
without a restart.
- Put is an upsert, so a controller can reconcile to it. Deleting a
provider or role that does not exist returns NotFound, as DeleteUser does
for a user; clients treat that as already deleted. The provider's account
ID travels in the request, since the filer does not know the STS
accountId.
- PutRole applies the IAM API's rules: AWS role names, at most 10 managed
policies.
- An S3 server serves the roles and providers of its own IAM config file
ahead of the store, so a stored entry with the same name has no effect
on that server.
- PutRole keeps a replaced role's RoleId and gives a role created anew a
fresh one, so sessions of a deleted role do not carry over to a later role
of the same name.
- DeletePolicy returns FailedPrecondition while a role attaches the policy
(see the IAM API's DeleteConflict in the previous change). DeletePolicy on
this service still does not check user attachments, which predates this.
- PutOIDCProvider requires an https issuer (http only for a loopback host):
STS fetches the issuer's signing keys from it, so over plain HTTP anyone
on the network path could substitute their own.
- The OIDC provider and role RPCs refuse to run on an unauthenticated
service (FailedPrecondition until jwt.filer_signing.key is set). Users and
policies keep the service's opt-in auth, but these grant STS access
outright: otherwise anyone who can reach the port could register an issuer
they control, create a role trusting it, and exchange a token for S3
credentials. The filer's unauthenticated notice becomes a warning that says
so.
- A store that cannot be read is Unavailable, never "not found", so a Put
never writes over an entry it could not see.
- Validation is shared with the IAM API through PrepareRoleDefinition and
PrepareOIDCProviderRecord.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* s3/iam: bind every role session to its role, and change roles atomically
Review follow-ups.
Session binding
- The role-ID check ran only when a session carried no policy names, and
AssumeRole embeds the role's attached policies, so those sessions kept
their permissions after the role was deleted or recreated. The check
now runs for every session carrying a role ID, before policy selection.
- A named role that cannot be resolved at issuance gets no session,
instead of one with no role ID (which nothing binds).
- A config-file role's ID is derived from its name and trust policy, not
the name alone: a different role put in the file under the same name
gets a new ID, while an unchanged role keeps its sessions across restarts.
Role writes
- RoleStore gains UpdateRole, a read-modify-write that lands only if the
role is unchanged since the read, and otherwise re-reads and retries. The
filer store uses the filer's write conditions (IF_NOT_EXISTS for a new
role, IF_ENTRY_EQUAL otherwise). CreateRole, UpdateAssumeRolePolicy and
Attach/DetachRolePolicy all go through it, so two gateways no longer
overwrite each other's changes, a change racing a delete no longer
writes the role back, and of two concurrent creates one gets
EntityAlreadyExists.
- The filer store's ListRoles pages past 1,000 entries and fails on a
broken stream instead of returning what arrived, so DeletePolicy's
attachment check sees every role. ListRoles skips a role deleted between
listing and reading it.
- CreateRole validates first; a failed write is ServiceFailure, not
InvalidInput. Any Tags.* parameter is refused, not only the first key.
- ExecuteAction's skipPersist covers the S3ApiConfiguration only; the
comment now says so. Role and OIDC provider actions write their own stores.
Each fix has a test that fails without it. Against a real filer with two
gateways, concurrent AttachRolePolicy calls lost 1-4 of 8 attachments per
run before this change and none after.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* iam: PutRole changes roles atomically and checks its ARN; https issuers' keys stay on https
Review follow-ups on top of the role-store changes.
- PutRole goes through RoleStore.UpdateRole, so the decision to keep an
existing role's ID or mint a new one is made against the role as it is
when written. A PutRole racing a DeleteRole can no longer write the
deleted role back with its old ID, which would revive its sessions. A
failed store read or write is Unavailable.
- PutRole refuses a role_arn that does not name the role: STS resolves a
role by the name in the ARN it is given.
- PutOIDCProvider requires an https issuer, but discovery could still name
a plain-http jwks_uri, and a key fetch could be redirected to http. For
an https issuer, a non-https jwks_uri from discovery is refused (the
issuer's own /.well-known/jwks.json is used instead), and the client
that fetches discovery and keys refuses any https-to-http redirect. An
operator-set jwksUri is left as configured.
Each has a test that fails without its guard.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* s3/iam: one role snapshot per decision; DeleteRole is atomic; watch a custom role store path
Review follow-ups.
- Authorization evaluates the policies of the role definition the session's
binding was checked against, instead of reading the role again: a role
replaced in between cannot lend a session its policies.
- AssumeRole and AssumeRoleWithLDAPIdentity issue the session from the
definition whose trust admits the caller (IAMManager.ResolveRoleForPrincipal),
and take its ID, duration cap and embedded policies from that same
definition. A role replaced after the caller's trust check by one that does
not trust the caller now yields AccessDenied, not a session bound to the
replacement.
- A RoleUpdate that returns nil deletes the role, on the same condition as a
write: the filer store deletes with ObjectTransaction on IF_ENTRY_EQUAL,
routed and locked like the conditional CreateEntry. DeleteRole decides
against the role it deletes, so a policy attached meanwhile on another
server is a DeleteConflict, and a delete never removes a role written
after its check.
- S3 servers watch the role store's configured basePath, not only
/etc/iam/roles, so a custom path also drops peers' cached roles on change.
Each has a test that fails without it. Live against a real filer: DeleteRole
refuses while a policy is attached and removes the entry once detached; all
test/s3/iam CI stages pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* s3/iam: state which roles DeletePolicy's attachment check can see
RolesAttachingPolicy sees the stored roles and this server's config-file
roles. A role defined only in another server's IAM config file is invisible
to it, so a config-file role that attaches a managed policy is protected
only on the servers whose file defines it. The doc comment now says so and
how to avoid it: keep such roles in every server's file, or attach only
config-file policies to config-file roles.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* iam: note that a role store set after startup is not watched for peer changes
S3 servers build their metadata watch list once, at startup, from the role
store installed then. SetRoleStore's doc now says that a filer-backed store
installed later with a different basePath is not watched, so peers' changes
to it reach this server's cached roles only when the cache expires.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* iam: DeleteRole deletes only the role it saw; issuer URLs are bare
Review follow-ups.
- The filer IAM service's DeleteRole looked the role up, then deleted by
name, so a PutRole landing in between had its new definition deleted. It
now deletes through RoleStore.UpdateRole, conditional on the entry it
read. If the role was replaced meanwhile, it returns Aborted rather than
deleting the replacement, and the caller decides again.
- PutOIDCProvider refuses an issuer URL with userinfo, a query or a
fragment. The provider's ARN comes from host and path alone, while STS
matches a token's iss claim against the stored URL exactly, so such a
provider shared the bare issuer's ARN and matched no token. A loopback
"localhost" is now matched without regard to case.
Both have tests that fail without them.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* iam: write OIDC providers atomically over the filer IAM gRPC service
PutOIDCProvider read the record, then stored unconditionally; a racing
DeleteOIDCProvider left the put's stale read merged into the rewritten
record. DeleteOIDCProvider read, then deleted unconditionally; a racing
PutOIDCProvider's newer record could be removed instead. These are the
races the role RPCs closed with UpdateRole.
OIDCProviderStore gains UpdateProvider with the same contract: memory
under its lock, filer as a conditional write (IF_ENTRY_EQUAL /
IF_NOT_EXISTS) or conditional delete retrying a changed entry.
PutOIDCProvider merges the fields the request cannot carry against the
record as it is written; DeleteOIDCProvider aborts rather than delete a
record replaced meanwhile.
isRoleWriteConflict is renamed isEntryWriteConflict — the conditional-
write check is shared by both stores now.
* iam: guard PutRole against a nil credential manager, fix its doc comment
PutRole read attached policies through s.credentialManager without the
nil check its sibling handlers make, so a server built without one
panicked on a PutRole naming a policy. It now fails the call as
FailedPrecondition like the others.
The doc comment also had the store/static precedence backwards: a stored
role shadows a same-named config-file role (as the overlay serves it),
not the other way around.
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
|
||
|
|
fa77cde7da |
vacuum: check compaction space against live bytes, not volume size (#11524)
* vacuum: size the compaction space check by live bytes, not volume size ensureCompactVolumeSpace required the volume's current .dat and .idx size as free space before compacting. That is the size of the garbage, not of what compaction writes, so on a disk that filled up until its volumes went read-only every compaction was refused, including all-garbage volumes that would compact to a superblock and an empty index. The sweep then retried every volume each cycle and reclaimed nothing (issue #11516). Estimate the output from what the needle map already tracks: live content bytes plus a per-needle framing upper bound behind a superblock, and one index entry per live needle. The estimate never exceeds the current volume size and preallocate still wins when larger. Volumes whose deleted sizes are unknown (.sdx converted back to .idx) keep the whole volume as the estimate. The disk probe moves behind a package variable so the tests can stand in for a full disk; the tests build real volumes instead of re-implementing the formula. * vacuum: space check reserves the index on top of preallocate, checks a separate index disk Review follow-ups: preallocate only stands in for the new .dat, so the rebuilt index is added on top of it; with separate index directories the data disk is checked for the .cpd and the index disk for the .cpx; and the estimates carry 1/16 headroom because counters rebuilt from an index file pass through a Bloom filter with a 0.1% false positive rate. Neither estimate exceeds the current file. * vacuum: split the space check by filesystem, not by directory name Two directories can sit on one filesystem and share its free space, so the data and index estimates are checked separately only when the index directory is on another device; otherwise the sum must fit. Unknown is treated as shared. * vacuum: ask the index directory for its share even when it looks like the same filesystem A volume mounted under the data directory's drive letter on Windows has the same volume name, so the identity check calls it shared. Checking the index directory for the index estimate as well costs one statfs and catches a full index mount either way. * vacuum: identify a Windows volume by its GUID, not its path prefix A volume can be reached through a drive letter and through a folder it is mounted on, so filepath.VolumeName says nothing about the free-space pool. Resolve each directory to its mount point and compare the volume GUIDs; when that fails the two are treated as shared. * vacuum: keep the framing and disk_space_low coverage the rebase displaced * rust volume: split the compaction space check across data and index disks Mirror the Go check: estimate the new .dat and rebuilt .idx separately — live content plus per-needle framing capped at the current file, with preallocate standing in for the data file when larger — and check each directory against its own filesystem's free space. Two directories on one filesystem are asked for the sum. * vacuum: tighten comments on the compaction space check Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
35b090a4df |
volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go) (#11554)
* volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go) An EC volume's deletion journal is a set of needle ids, but shard copy and index recovery appended the peer's whole journal, doubling the file on every ec_balance round trip. Fold the peer's ids in as a union instead: only ids the local journal lacks are appended. - The journal is never replaced. A mounted EcVolume merges a peer's ids through its live handle under the lock deletes take (Go MergeJournal / Rust merge_journal), wherever its journal lives. - An unmounted journal gets only the missing ids appended while mounts are excluded; the delta is read outside the lock and re-read if the journal changed. - The source .ecj streams into memory as an id set: no staging files, chunked reads, memory proportional to distinct ids. - Go and Rust agree that a source journal exists when it sends a modified time or any bytes. A missing source stays a no-op. - Rust runs every merge in spawn_blocking and shares one receive/merge path between shard copy and index recovery. The decode path and the journal format are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: route .ecj merges to the runtime that holds the journal open Disks sharing one index directory all resolved as the journal's owner, so the last one won and a sibling's mounted runtime was skipped: the merge appended behind its open handle and the sibling kept serving the peer's deleted needles until remount. Callers now name the receiving disk by its data directory; the merge goes through that disk's runtime, else a sibling runtime whose journal is the target file. In Go the unmounted append now holds every disk's EC lock (in location order) while it rechecks for a mount, so a sibling mounting from this disk's index during the unlocked read is merged through instead. In Rust a mount that lands during the read is merged through directly and its added count returned, rather than discarded and reported as zero. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: sync merged .ecj records outside the disks' EC locks The unmounted merge held every disk's EC read lock across its fsync, so a slow sync on one disk held off mounts on all of them, along with the EC reads queued behind those mounts. Mounts only need to be excluded while the records are written: the write now happens under the locks and the fsync after they are released, since a later mount reads the written records from the page cache. A failed fsync rolls back only if nothing has mounted the journal or appended to it since the write. A merge through a mounted volume now keeps only that volume's disk locked across its fsync. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: roll back an unsynced .ecj merge through a volume mounted mid-sync If a volume mounted after the unmounted merge wrote its records but before the fsync failed, the rollback kept the records because the journal was now open, leaving ids in the volume's deleted set that may never reach disk; a retried merge then saw them and synced nothing. The rollback now goes through that volume the way its own failed journal fsync does: truncate back and drop the ids from the in-memory set, so a retry appends and syncs them again. It still keeps the records if the volume journaled since, as truncating would lose that delete. No fsync runs under the disk locks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: decide .ecj merge rollback from the journal's actual length Two runtimes can hold one journal (cross-disk mounts). The rollback of an unsynced merge checked one runtime's cached ecjFileSize, which another runtime's appends leave stale, so it could truncate a delete that runtime had already synced. The rollback now holds every holder's journal lock and truncates only if the file's actual length is still the append's end, then updates each holder's size and deleted set. Otherwise later records follow the merged ones, so they stay and are rewritten in place and synced outside the locks, rather than left possibly not durable. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: keep unsynced .ecj merge ids out of mounted deleted sets When a merge's fsync failed, later records blocked the rollback, and the rewrite-and-sync failed as well, the merged ids stayed in every mounted volume's deleted set without being shown durable, so a retried merge saw them as present and synced nothing. They now leave those sets while the records stay in the file, matching DeleteNeedleFromEcx, which publishes an id only after its record syncs. The merge returns the error and a retry appends and syncs them again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: publish merged .ecj ids to every holder of the journal Two runtimes can journal into the same file when disks share an index directory. The merge went through only the first holder, leaving a sibling's in-memory deleted set without the ids, so it could keep serving a needle the peer deleted until it remounted. Every holder of the journal now gets the merged ids, in Go and in the volume server. Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Publish merged .ecj ids to the journal actually written mountedEcJournal prefers the receiving disk's own runtime for the vid, whose journal may live in its data directory while the copied records name a sibling's journal in the index directory. Publishing by the requested ecjPath then marked a holder of a different file deleted on records that file never persisted, resurrecting the needles on remount. Publish by the picked runtime's journal path instead. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
1445960f8c |
filer: persist pending chunk deletions across restarts (durable deletion ledger) (#11550)
* fix(filer): persist pending chunk deletions across restarts The in-memory FileIdDeletionQueue and DeletionRetryQueue lose every queued-but-unconfirmed deletion when the filer process restarts. Because deletions only enter the pipeline through that queue, a crash between enqueue and the volume confirming the delete leaks the chunk permanently: nothing remembers it. In a multi-filer deployment this was observed as growing collections of orphaned chunks after filer restarts, and — via meta-replay from a peer that still had the entry — orphans being "resurrected" as live references on the recovered filer. This implements the "periodic snapshot with recovery on startup" option noted in the existing DeletionRetryQueue TODO, using the store's KV layer (no new iterator API required across the 15+ store backends): - queueDeletions() is the single entry point that keeps the hot in-memory queue and the durable ledger in sync. - Only terminal outcomes (success / not-found / permanent) remove an id from the ledger; retryable failures keep it, which is the point. - A timer and Shutdown() snapshot the pending set to a single KV key. - On startup, reloadDeletionLedger() re-queues recovered ids after a grace window so the initial peer meta-aggregation settles first. This avoids a new hazard: purging a chunk that a lagging peer is about to re-reference as live data (stale replay turns a stale read into a dangling read otherwise). - Volume deletes are idempotent (not-found == success), so re-deleting after a crash never double-frees. - Kill switch via viper: filer.deleteQueue.persist=false opts out entirely (reload also refuses to recover so a stale ledger never comes back). Tunables: filer.deleteQueue.persistInterval, .recoveryGrace. Adds unit tests covering snapshot+recover, retry-keeps-entry, disabled switch, and zero-value Filer safety (run green under -race). Co-Authored-By: Athena 🏛️ <hermes-agent@local> (custom / Qwen3.8-Flash-Next-ROCmFP4) * filer: harden the deletion ledger - Scope the ledger key by filer address so filers sharing one store do not overwrite each other's pending sets; ledgers written under the old unscoped key are claimed once on startup. - Serialize snapshots on deletionSnapshotLock so an in-flight timer snapshot cannot overwrite a newer shutdown snapshot, and wake the snapshotter on every queue/forget so a queued id persists within milliseconds instead of a full interval. - Merge recovered ids into the pending set immediately on reload; only the queue push waits out the grace window, so an early snapshot rewrites the recovered ids rather than dropping them. - A failed or unparseable ledger read blocks persistence for the run instead of letting snapshots overwrite the unread ledger. - Split the ledger into part keys when it exceeds one 64KB value so stores with a size cap (FoundationDB) do not strand the backlog. - GetReadyItems reports retry-exhausted ids so they are forgotten in the ledger instead of replaying after every restart. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: close the remaining deletion-ledger durability gaps - A manifest referencing a missing part is corruption: surface a wrapped error and block persistence instead of treating the ledger as absent. - Multipart snapshots write generation-scoped part keys and publish the manifest last, so a crash never mixes old and new part contents. - Orphaned parts are tracked in a persisted .stale sidecar and retried. - Legacy/index ledgers are republished under the scoped key before the old keys are removed. - A ledger index lets a filer restart under a new address claim the ledger its previous incarnation left behind. - Expired and permanently-failed retry items only forget the ledger epoch they recorded, so they cannot erase a re-queued id. - A failed startup read no longer disables persistence: every snapshot retries the reload until the store reads again. * filer: tighten ledger claiming, index updates, and retry epochs - touchLedgerIndex verifies its write and retries so a concurrent filer's merge cannot silently drop this key from the index. - Foreign-ledger claims abort on any unreadable source instead of leaving it stranded once the new scoped key exists. - A source that republished during the claim is left in place and its newer ids merge into the claimant's pending set. - AddOrUpdate no longer overwrites the ledger epoch of an in-flight retry item, so its expiry or permanent outcome cannot forget a record that was re-queued after the attempt began. - The recovery grace wait exits on shutdown instead of re-queueing after the filer has stopped. * filer: requeue surviving records, persist claim deltas, guard index writes - A dropped retry item (expired or permanent) whose ledger record was re-enqueued now pushes the id back through the hot queue instead of leaving it pending with nothing scheduled. - Ids merged from a claim source that republished mid-claim are rewritten under our ledger immediately, so they are durable even if the claimant crashes before the next snapshot. - touchLedgerIndex aborts when the index read fails for a real error; only ErrKvNotFound means the index is empty, so a transient failure can no longer wipe peer entries with a one-key write. --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
43abe21ffa |
Helm: Add Opt-in Read-Only Root Filesystem Support (#11562)
* security context changes Signed-off-by: Subhadeep Maity <smaity@slb.com> * root file changes Signed-off-by: Subhadeep Maity <smaity@slb.com> * Skip tmp mount when extras provide one; use allInOne context for the bucket hook * Mount tmp for secondary containers; keep user /tmp on the main container --------- Signed-off-by: Subhadeep Maity <smaity@slb.com> Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> Co-authored-by: Subhadeep Maity <smaity@slb.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
67034dee12 | docs: regenerate star history chart | ||
|
|
8d97284d0a |
filer: option to store system metadata logs in their own collection (#11551)
* feat(filer): option to store system metadata logs in their own collection The filer's internal /topics/.system/log chunks are assigned to the filer's default collection (-collection). In a multi-filer deployment that default is often empty, so every restart flap, full-sync, or event-buffered flush grows the default collection with system chunks that are indistinguishable from user data in collection.list. This is a large part of what makes the default collection balloon and confuses orphan analysis. This keeps the internal log in a dedicated collection when the operator asks for one, without changing where user data goes: - New optional override, filer.options.metaLog.collection (and .replication), read in NewFiler so both `weed filer` and `weed server -filer` honour it. Default "" => exactly today's behaviour (log follows the filer default), fully backward compatible. - Resolution is a small helper: override first, then the filer default, then a storage rule matched on the log path. Kept separate from the user write path so the internal log targets itself. - bucketCollection() is hardened the same way it already protects the filer's default collection: a bucket that happens to resolve to the redirected meta-log collection must not drop it on delete, because it backs internal log volumes. - Scaffold filer.toml documents the new knobs under [filer.options]. Related to the persisted deletion ledger branch (fix/persist-deletion-queue): together they cut the two sources of post-flap junk in the default collection — that PR stops orphaned user-chunk leak on filer crash, this one stops the internal log from living in default at all. They are independent: no file overlap, no functional dependency; either can merge first. They are paired only in the narrative of cleaning up default. Adds unit tests for the collection/replication resolution chain, the viper keys, and the bucket-delete guard (run green under -race). Co-Authored-By: Athena 🏛️ <hermes-agent@local> (custom / Qwen3.8-Flash-Next-ROCmFP4) * filer: collect bucket chunks when its collection survives the delete bucketCollection returning "" preserves the collection, but the bucket path still skipped per-entry chunk collection and could skip listing the children entirely, so a bucket sharing the meta-log (or any preserved) collection left its object chunks orphaned with no entry pointing at them. Only the wholesale drop of a deleted collection skips those now. Note in filer.toml that the meta-log target should stay stable: chunks written under an older collection are not migrated. * filer: exercise the metaLog override wiring through NewFiler The viper test only echoed back the keys it set, so a wrong key in NewFiler would still pass. It now asserts the fields NewFiler fills from those keys. Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: tighten comments around the metaLog collection override Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Chris Lu <chris.lu@gmail.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
d8f926cf46 |
volume server: stream needle chunks without the store lock (#11488)
* volume server: read GET/HEAD needles off the store lock, and only once The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stream needle chunks without the store lock StreamingBody::poll_frame took store.read() and find_volume for every chunk to compare the volume's compaction revision, dup'd the source handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the stream also holds a data-file read lease for its whole life, while a writer waits for that lease under store.write(): the next chunk's store.read() then waits for the writer and the writer for the stream. The per-chunk re-lookup was also wrong. The stream reads a handle opened at plan time, which pins the .dat inode the offset was resolved against; a vacuum commit renames a new file over .dat and leaves that inode untouched. The re-looked-up offset belongs to the new file but was read from the old inode, so a stream whose needle a vacuum moved ended in a checksum error. The pinned offset stays valid, so the check, and with it every store access, is dropped, along with the now unused re_lookup_needle_data_offset and the revision fields of the read plan. The source is shared as an Arc instead of dup'd per chunk, and the chunk buffer is a BytesMut that the blocking read hands back with its result, so its allocation is reclaimed once the previous frame has been written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a needle stream once its volume becomes unavailable Taking the store lock out of StreamingBody also dropped its per-chunk unavailable_error() check. With -hasSlowRead a writer can take the data-file lease between chunks, fail its fsync and its truncate, and mark the volume unavailable; the stream then kept serving the rest of the needle from its pinned handle. The volume's io_unavailable reason is now an Arc-shared leaf mutex that the read plan hands to the stream. Each chunk checks it under its data-file lease, where the writer marks it, and fails with the same "volume is unavailable: <reason>" error the old check returned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
8a9563e53d |
volume: test that only the disk holding the replaced replica gets the free-slot credit (#11486)
* volume server: split volume_copy into phases and type the delete-after-status gate volume_copy was one ~400-line handler, and the rule that an existing local replica is deleted only after the source's ReadVolumeFileStatus succeeded was held by statement order alone. The keep_remote_data=true that the pre-copy delete and the failed-copy rollback must share was kept in sync by a comment pointing from one to the other. The handler is now a ~60-line orchestrator over connect_to_copy_source, SourceVolumeStatus::fetch, delete_existing_replica, plan_copy_destination and a VolumeCopyJob whose run() drives preallocate_dat, transfer_files, finish_copied_files and mount_and_reply, with cleanup_failed_copy on error. delete_existing_replica takes a &SourceVolumeStatus, which only fetch can construct (private field in a child module), so the delete cannot be called before the status RPC. Both deletes go through delete_replica_keep_remote. Pure refactor: call order, status codes and messages, cancellation checks, throttling, progress reports and cleanup are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: find space for a VolumeCopy before deleting the replica it replaces VolumeCopy deleted an existing local replica as soon as the source answered ReadVolumeFileStatus and only then looked for a location with room for the copy. With no usable location (disk full, low-disk, wrong disk type) the call errored after the delete, leaving the node with neither the old replica nor the new one. Plan the destination first, as Go does: find_free_location_replacing credits the location holding the replaced volume with that volume's slot, so a disk at its volume limit that holds the replica still accepts the copy. Only then delete the replica and write the .note (still after the delete, as in Go). delete_existing_replica now takes the planned CopyDestination, so the delete cannot precede the plan. find_free_location_predicate keeps its behaviour. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: describe the replace-credit test against the current VolumeCopy flow Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
90f8c4378f |
s3: restrict admin gRPC to local callers when no signing key (#11530)
* s3: restrict admin gRPC to local callers when no signing key The S3 gateway's gRPC port (default 0.0.0.0:19000, always on) serves the IAM cache and internal lifecycle admin services. checkAdminAuth was a no-op when jwt.filer_signing.key was unset, so any reachable host could PutIdentity an admin identity and take over the bucket data. Without a shared key callers cannot be distinguished, so admin RPCs are now limited to unix-socket, loopback, and the server's own interface addresses. Remote filer-to-S3 propagation and lifecycle workers must set jwt.filer_signing.key; the Bearer-token path is unchanged. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: fail closed on nil guard and refresh local addresses per call Review feedback: a nil filerGuard bypassed all checks — treat it like a missing key and require a local peer. The own-address set was cached forever, so interfaces added later were rejected; enumerate per call instead since admin RPCs are rare. Nil ctx is denied rather than panics. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: read the signing key once and bound interface enumeration Review feedback: reading SigningKey twice could straddle a SIGHUP reload — an old nonempty key skipped the local-peer check while the new empty key verified the token. And enumerating interfaces per no-key call is wasteful for co-located workers dialing the announced address; cache the address set for 30s so new interfaces still become usable promptly. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: enumerate interface addresses per no-key admin call A cached address set keeps trusting an IP after it is removed from the host and reassigned to another machine — that host would then hold unauthenticated admin access for the cache TTL. Per-call enumeration only runs for non-loopback TCP peers on the no-key path, which is low-volume admin traffic, so the freshness is worth the syscall. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
eeec9ec09a |
volume server: refuse to compact a volume tiered to remote storage (#11545)
* volume server: refuse a tier move while compacting, and a commit once tiered A tier move to remote and a vacuum compaction of the same volume could interleave and leave the volume unreadable: - A compaction committing while the upload ran swapped .dat/.idx under the transfer, which reopens the .dat by path per part. The move then published an object holding the old (or a mixed) layout against the compacted .idx, and with keep_local_dat_file=false deleted the only compacted .dat. - A tier move finishing while the compaction copy ran (or between the copy and the commit) let the commit swap in the compacted .idx while the reload served the pre-compaction remote object through it. The tier move now refuses to start while the volume is compacting, and re-checks the compaction revision under the store write lock before it records the remote file; on a mismatch it deletes the uploaded object and fails with FailedPrecondition, leaving the volume local. Committing a compaction on a volume that has a remote file is refused and its .cpd/.cpx removed, since the reload would read the remote object through the compacted index. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: abort a tier move whose volume was replaced or removed The tier-up bookkeeping looked the volume up by id only and compared the compaction revision. A delete and re-create of the same id during the upload yields a fresh volume at the same revision, so the move recorded the old volume's object on the new one and, without keep_local_dat_file, removed the new .dat. An unmounted volume was skipped and the move reported success, leaving the uploaded object referenced by nothing. Capture the volume instance (its data-file access control Arc, as the scan and read plans do) with the revision, and require both under the store write lock. A replaced volume fails with FailedPrecondition, a missing one with NotFound; either way nothing is recorded and the object is deleted after the lock is released. Go fails in both cases because deleting or unmounting closes the descriptor its copy reads. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: refuse to compact a volume tiered to remote storage Committing a compaction of a tiered volume is refused, since the reload would read the remote object through the compacted index. The compaction itself still started: a tiered volume's data backend is the remote object (the local .dat is dropped or deleted on tier-up), so an explicit vacuum streamed the whole .dat out of remote storage into a .cpd that the commit then discarded. Refuse at the start of the compaction instead, before the .cpd is created, at the point where Go's copy opens the local .dat. The truncated-index test now uses a read-only local volume for its sorted index, since a tiered one no longer reaches the copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: match the ReadOnly(VolumeId) variant in write_volume_needles #11543 matched VolumeError::ReadOnly as a unit variant in Store::write_volume_needles, and #11544 changed it to ReadOnly(VolumeId) in the same merge window. Each passed CI on its own, but master no longer compiles the Rust volume server. Carry the volume id through. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
7c7834c98e |
Helm: Add Configurable Security Contexts for Chart-Managed Workloads (#11558)
* security context changes Signed-off-by: Subhadeep Maity <smaity@slb.com> * extending examples Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> * updating examples Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> * examples Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> * Update k8s/charts/seaweedfs/values.yaml Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> * Update k8s/charts/seaweedfs/values.yaml Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * review comments Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> --------- Signed-off-by: Subhadeep Maity <smaity@slb.com> Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com> Co-authored-by: Subhadeep Maity <smaity@slb.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> |
||
|
|
2a42d56437 |
ecbalancer: honour total-shards-per-rack cap in Place / PlaceDurabilityFirst (#11553)
* ecbalancer: honour total-shards-per-rack cap in Place / PlaceDurabilityFirst
Worker auto-EC encode places via Topology.Place, which capped each shard
type independently (ceil(data/racks), ceil(parity/racks)). On an 8-rack
topology that permits 3 total shards on one rack, so losing two racks
strands 6/14 and a 10+4 volume becomes unreadable.
- tryPlace caps the total shards (data + parity) per rack in both modes,
whether or not ReplicaPlacement is set.
- rackTotalCap picks the smallest per-rack total the racks' real room
(free slots, bounded by the per-disk cap and node free slots, counting
shards already placed) can satisfy. On a uniform cluster it is
ceil(shards/racks); a nearly full rack raises it just enough that the
cap alone never fails an encode.
- PlaceDurabilityFirst gets a last rung that drops the rack cap
("rack-total-cap" in Relaxed), so it fails only when no disk has room.
PlaceStrict keeps the cap as a hard limit.
- chooseShardDest tries the next rack when the chosen one has no node
that fits, and room checks count the per-disk cap, so a rack whose
disks are all at the cap is no longer picked and then failed on
(pre-existing: 3-node rack + single-disk rack failed at shard 9).
- Docs no longer claim the cap guarantees surviving rack loss; the
placement error names the caps in effect; the encode warning no longer
says replica placement when other constraints were relaxed.
place_rack_cap_test.go covers 10+4 over 8 racks (max 2/rack, 3/rack on
master), a starved rack, nearly full racks, the preferred-tag tier, the
full-disk rack, and rackTotalCap directly.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ecbalancer: size the rack total cap from room left under SameRackCount
The rack total cap counted each rack's free disk room, but attempts that
enforce ReplicaPlacement also stop a node at SameRackCount shards. With
SameRackCount=1, four one-node racks and four three-node racks got cap 2,
which fits only 12 of 14 shards: strict placement failed and
durability-first relaxed replica placement although 1 per small rack and
up to 3 per large rack fits.
Attempts that enforce ReplicaPlacement now use a cap sized from each
node's remaining SameRackCount allowance; attempts that relax it keep the
disk-room cap.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|