Commit Graph
10048 Commits
Author SHA1 Message Date
Chris LuandDevin 4bf5e92b64 storage: unify needle-map counters across index loaders and warn on copy-count drift (#11701)
* storage: count index-load metrics only for live-key removals

The in-memory and leveldb offset loaders counted a deletion for every
tombstone row replayed, including re-deletes of already-deleted keys
(runtime logDelete only counts a live removal), and added the returned
old size to the byte counters even when it was a tombstone sentinel —
uint64(-1) wraps DeletionByteCounter. The leveldb offset loader also
counted every index row as a file and its bytes unconditionally. Gate
the counters the same way the runtime paths do so a reload reports the
same numbers the runtime counters hold.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* storage: derive index metrics from latest-row state in the metric loader

The leveldb metric loader walked the index in reverse and incremented
FileCounter for every row (inflating FileCount with tombstones) and
DeletionCounter for every superseded row plus every deleted-key row
(double counting both). Per key, the runtime deleted count is exactly
(valid rows) - (1 when the latest row is live), so count live-latest
rows via the same bloom filter and derive both deletion counters; this
matches the runtime logPut/logDelete counters exactly instead of
drifting with tombstone history.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* server: warn instead of failing volume copies on counter drift

checkCopyCounts compared the source's in-memory counters against the
target's fresh-load counters, but counters computed by different
loaders drifted apart for identical .idx data, so a good copy was
rejected (issue #11678). The file-size checks already gate integrity;
log the mismatch instead of failing the copy.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* storage: assert MaxFileKey parity in the metric loader test

MaxFileKey derives from row order alone, so unlike the bloom-filtered
deletion counters it must always match the runtime counters exactly.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-10 22:31:46 +08:00
Peter Dodd 5f1a742726 fix: return NoSuchKey and drop stale entries when a remote-mounted object is gone from the remote (#11680)
A remote-only entry whose object was deleted from the remote storage
outside the filer answered GET with 500 and stayed in the filer. The
remote's not-found was lost on the way: the backends' ReadFile returned
it as an untyped error, so FetchAndWriteNeedle failed with codes.Unknown
and nothing downstream could tell it from any other failure.

- remote_storage: GCS, S3 (NoSuchKey) and Azure (BlobNotFound) reads
  return ErrRemoteObjectNotFound. GCS reports a missing bucket the same
  way as a missing object, so it confirms the bucket with a listing.
- volume server and filer: the not-found crosses gRPC as codes.NotFound
  carrying the sentinel's text, and the filer's cache RPC returns
  codes.NotFound, which the S3 gateway already maps to NoSuchKey.
- s3api: the origin fallback answers NoSuchKey on a confirmed not-found.
- filer: a confirmed not-found removes the stale entry, so the lazy
  remote-metadata cache converges on the remote. Only remote-only files
  outside .versions and without an active object lock are removed, only
  if unchanged since the fetch (checked on the object's write owner,
  under the lock S3 object writes take), and with a metadata-only
  delete: the filer skips its inline remote delete, and the delete
  events' entries carry a marker that makes filer.remote.sync and
  filer.remote.gateway skip their remote delete, while filer.sync still
  replicates it. The replicated DeleteEntryRequest carries
  keep_remote_object, so the destination's delete events are marked too.
  The store drops the marker from every write, so clients cannot plant
  it.
2026-10-10 11:00:04 +08:00
d3dc03c85a s3api: resolve volume data encryption in CopyObject SSE flows (#11646) (#11683)
* s3api: resolve volume data encryption in CopyObject SSE flows (#11646)

* s3api: honor bucket-default KMS key on copy and fix transformed-upload metadata

- Synthesize the destination bucket's default encryption as request
  headers before any SSE evaluation, so the configured KMS key ID and
  bucket-key setting reach the copy paths instead of only a boolean.
- uploadTransformedChunkData returns the upload result so callers record
  the uploader's cipher key AND compression decision; a wrongly cleared
  IsCompressed made transformed copies unreadable.
- decompressChunkVolumeCipher fails loudly when a compressed chunk does
  not decompress, instead of uploading still-compressed bytes marked
  uncompressed.
- copyMultipartSSECChunk now strips the volume cipher and re-encrypts on
  upload like the other transform paths.
- The ciphered inner upload now uses the caller's private BytesBuffer.

* s3api: extract copy bucket-default header synthesis for coverage

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: test bucket-default encryption header synthesis on copy

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: validate copy encryption headers before bucket defaults and resolve empty KMS key

* s3api: name the AWS-managed SSE-KMS default key once

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-10 10:59:36 +08:00
ea65746947 fix(filer): end metadata subscriptions before gRPC GracefulStop on shutdown (#11663)
* fix(filer): end metadata subscriptions before gRPC GracefulStop on shutdown

GracefulStop waits for every open stream. The filer's own MetaAggregator
subscription and, under -s3, the S3 gateway's IAM subscription live in
the same process and only end when the filer shuts down, which happens
after GracefulStop - so every shutdown sat out the full 15s timeout.

StopSubscriptions ends them (and any started later) right after leaving
the lock ring; the handlers derive their context from the stream and
this signal. Measured on weed server -s3: stop 25.6s -> 10.7s, and
0.6s with -volume.preStopSeconds=0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FTq6bDpgdQfqQagQsUUvdw

* fix(filer): end stopped subscriptions with Unavailable, not cleanly

A subscription that StopSubscriptions cut short returned nil. The client
reads that as io.EOF, "caught up, done", and util.RetryUntil - which the
S3 gateway and mount follow with - stops on nil. A separate S3 gateway
therefore never resubscribed after its filer restarted (measured: 0
subscriptions in 25s after the restart; upstream master and this fix: it
is back within a second).

endOfSubscription now turns that clean end into codes.Unavailable, which
is what a dropped connection looks like, so followers reconnect. Errors
pass through, and so does the end of a stream the client closed.

A subscriber that has stopped reading is still left to GracefulStop's
timeout: its handler sits in stream.Send, which only the end of the
stream releases, and grpc-go does not allow Send after the handler
returns. Documented at StopSubscriptions. Shutdown under weed server -s3
is unchanged at 10.7s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FTq6bDpgdQfqQagQsUUvdw

* fix(filer): interrupt in-flight subscription reads and end them Unavailable

StopSubscriptions only reached the wait points: a handler replaying the
ring or the persisted log kept streaming until the pass ended, and a pass
cut short at a ctx.Err() boundary (the replay semaphore, chunk ref
collection, ref sends) returned a wrapped context.Canceled, which reads as
codes.Canceled on the wire - a code transient-error classifiers treat as
non-retryable.

eachLogEntryFn and sendRefsBatched now check the subscription context per
entry/batch, and endOfSubscription reports a context.Canceled cut short by
the stop signal as codes.Unavailable, same as a clean end it interrupted.

* filer: honor the subscription context in the pipelined sender

A subscriber that stops reading wedges sendLoop in stream.Send; Send and
Close then blocked past StopSubscriptions, so GracefulStop still waited out
its timeout. Both now return when the subscription context ends, letting the
handler return and gRPC tear the stream down.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: stop sendLoop from starting sends after the subscription ends

With Send and Close released by the subscription context, sendLoop could
still issue a stream.Send after the handler returned. Gate every send on
the context so at most the in-flight Send overlaps teardown - and the
stream's own teardown is what unblocks it.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: assert sendLoop exits in the cancel test

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-10 10:59:00 +08:00
Chris Lu c023493cf0 s3api: tolerate session tokens on statically configured credentials (#11694)
* s3api: consolidate session token extraction into extractSessionToken

Three call sites duplicated the same header/header/query lookup; share
one helper named after the existing s3tables equivalent.

* s3api: tolerate session tokens on statically configured credentials

Credential vendors like Unity Catalog emit a session token with every vended credential, including static ones (UC's StaticAwsCredentialGenerator only engages when s3.sessionToken is set). Requests signed by a configured access key were routed to STS validation and rejected, so static credential vending never worked against SeaweedFS.

Resolve the access key first: when it maps to a configured credential the signature alone authenticates the request, and the attached token is marked ignored so authorization does not route it into the STS session-policy path. STS-issued access keys are never in the static map, so temporary credentials still validate their token exactly as before.

* s3api: cover static credentials carrying a session token

* test: exercise UC static credential vending against SeaweedFS

s3.sessionToken.0 selects UC's StaticAwsCredentialGenerator, which vends the configured keys verbatim. The vended session token is foreign to SeaweedFS and previously failed SigV4; now it round-trips through temporary-table-credentials into real S3 I/O.

* s3api: reject temporary credentials in GetFederationToken after auth

Token presence alone cannot distinguish a temporary credential from a
statically configured one carrying a vended token. Move the check behind
verifyV4Signature and key it on the operative session token so tolerated
tokens keep the caller eligible.

* s3api: test GetFederationToken with authenticated temporary credentials

Sign the rejection cases with real session credentials so they reach the
post-auth check, and cover a vended static credential being accepted.

* test: check vended-credential delete error in UC integration test
2026-10-10 10:11:24 +08:00
Paolo Valletta 4d306e1279 s3api: revoke the identities a config reload no longer declares (#11679)
* s3api: revoke the identities a config reload no longer declares

A static config file reload merges with isFullState=false, so an identity
deleted from -s3.config kept authenticating with all of its access keys until
the process restarted. An operator revoking a leaked key got no error and no
revocation.

The file is the source of truth for the identities it declares, so a reload now
drops the ones that have left it, together with their access keys and those of
their service accounts. staticIdentityNames was only ever added to, which kept a
removed name protected as well; the names the file itself declares are tracked
separately from the AWS environment credentials, so a reload never revokes what
the file never declared.

* s3api: keep a file reload from replacing the dynamic store

Removing the last identity from a config file emptied staticIdentityNames, so
useStaticConfig turned false and the next reload of that file took the replace
path: every filer-managed identity and its access keys disappeared until a
dynamic reload brought them back.

A static config file is authoritative for the identities it declares and never
for the dynamic store, so a file load now always merges. The first file load at
startup takes the merge path as well, from an empty state.

Reported by the Devin and Greptile reviews of this PR.
TestReloadStaticConfigWithoutIdentitiesKeepsDynamic reloads an emptied file twice
and asserts that a filer-managed identity and its access key survive both; the
existing test now also checks that the service account key was loaded before the
reload and that the environment identity's key still works after it.

* s3api: clear the environment credentials in the emptied-file test

The AWS environment identity is static and is re-added after every merge, so on
a runner that has AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY set it kept
hasStaticConfig true once the file was emptied, and the next reload still took
the merge path: the regression this test guards passed unnoticed.

Clear both variables whatever the runner has set. With the pre-fix condition
(hasStaticConfig alone) and the variables present in the environment, the test
now fails with "reload 2: a filer-managed identity must survive a reload of an
emptied file".

Reported by the Greptile review of this PR.
2026-10-10 09:34:11 +08:00
Chris LuandDevin 5d80ac39b8 s3api: verify under the write lock before re-committing an ambiguous routed PUT (#11649)
* s3api: verify under the write lock before re-committing an ambiguous routed PUT

A routed PUT whose response was lost after the owner committed, or whose
transaction returned a store error, fell straight into the lock path and
re-sent the same entry. A concurrent PUT that had superseded the commit in
the meantime had already deleted this entry's chunks as old, so the
re-commit restored metadata pointing at dead needles and the object read
back 404 permanently.

The lock path now resolves the routed attempt's outcome inside the write
lock first: a stored entry with the uploaded chunks means the route
landed; a different stored entry or an unresolved lookup refuses the
re-commit with ServiceUnavailable so the client retries with a fresh
upload; only a proven-absent entry falls through to the normal create.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: tighten ambiguous routed PUT recovery

- a match found only after an unanswered filer is not authoritative;
  the unreachable filer may hold a newer entry, so refuse instead of
  finalizing on it (both recovery paths)
- a failed post-recovery finalization now rolls the recovered entry
  back, matching the create path's undo
- a proven-absent entry is only re-committed after probing that the
  uploaded chunks' needles are still alive; a committed-and-deleted
  PUT would otherwise write back an entry pointing at reclaimed
  needles
- the .versions latest pointer no longer flips back to an older
  version when a newer one committed while the write was uncertain

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: stamp late versions noncurrent when a newer latest pointer wins

The keep-newer early return in updateLatestVersionInDirectory left the
version just stored without ExtNoncurrentSinceNsKey, so the lifecycle
engine could never age it out.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: only a failure past mutation 0 makes a routed PUT ambiguous

A deterministic refusal at the PUT mutation (e.g. "existing entry is a
directory") applied nothing, but marking it ambiguous sent the lock-path
fallback through conservative recovery, which found the directory entry
and refused with 503 instead of the correct 409.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: keep all transaction errors ambiguous; route directory conflicts through the lock path

* s3: verify uploaded chunks before re-committing over a directory

A stored directory does not prove the routed PUT never committed: the
route can land, a delete can remove the entry and its chunks, and a
nested write can recreate the directory before recovery takes the lock.
Re-committing then stores an entry pointing at dead needles. Probe the
uploaded chunks first and refuse with ServiceUnavailable when they can
no longer be verified.

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-10 08:52:34 +08:00
Chris Lu e337443176 filer: cut ReaderCache mutex contention on small reads (#11677)
* filer: skip cache lock when a cacher cannot be removed yet

SingleChunkCacher.readChunkAt defers removeConsumed on every read, and
every unpin retries removal too. Each call acquired the ReaderCache
mutex even though the removal conditions are plain atomics, so a busy
cache paid a lock acquisition per read just to discover readers>0.

Check the atomics before locking: when a cacher is not yet consumable
the removal is impossible and the lock round trip is pure contention.
Removals still run under the lock via removeConsumedLocked, so the
attach-vs-remove race keeps its existing serialization.

Ref #11676

* filer: guard stream position with a per-stream mutex, not the cache lock

chunkStream.cacher is only ever shared by concurrent ReadAt calls on one
ChunkReadAt, yet pin, unpin, releaseIfFinished, and releaseStream all
mutated it under the ReaderCache mutex that serializes every reader in
the process. Each read that switched chunks took that global lock to pin
the new chunk, released it, then reacquired it in unpin() just to drop
the old chunk's pin and retry removal; releaseIfFinished and
releaseStream did the same lock-detach-unlock-relock dance. At high
small-GET concurrency that turned per-request stream bookkeeping into a
global convoy (#11676).

Give chunkStream its own mutex and drop the cache lock from the pin
lifecycle entirely: pin/detach serialize on stream.mu, the pin counter
and consumable checks are already atomics, and removeConsumed only
takes the cache lock when a cacher is actually removable. The map lock
now guards only map membership and read registration.

* filer: look up cached chunks under the read lock

readChunkAt serialized every small read on the write lock even though
the common paths are read-only: an existing downloader just needs its
read registered, and a chunkCache hit needs no map access at all. Each
GET also paid the lock a second time to reach the chunkCache check.

Use the read lock for the downloader lookup and registration; the
registered read keeps the buffer alive against a concurrent destroy,
and only error eviction, insertion, and removal need the write lock.
The chunkCache probe runs lock-free, and the insert path re-checks the
map under the write lock to cover a downloader registered in between.

* filer: start the chunk download outside the map lock

The insert path held the ReaderCache write lock across goroutine spawn
and the cacheStartedCh handshake, so every downloader miss serialized
against the startup of a fetch goroutine. Register the cacher in the
map under the lock, then start the download after releasing it; a fetch
that fails early still lands in the map and is evicted by the next
reader's completed-error check.

* filer: test that stream pin lifecycle stays off the cache lock

Regression coverage for the contention fix: pin and releaseStream on a
non-removable cacher must complete while the ReaderCache lock is held
by another goroutine.

* filer: pin the stream under the map lock

Between read registration and stream.pin the cacher showed zero pins,
so a budget eviction in that gap could pick a chunk the stream was just
attaching to and the stream's next slice refetched it. The pin counter
is an atomic and stream.mu is never held while acquiring the cache
lock, so pinning inside the map hold is deadlock-free and closes the
window.
2026-10-10 06:52:37 +08:00
838c554e33 s3api: reject SSE headers PutObject cannot honor, as CopyObject does (#11642)
* s3api: reject SSE headers PutObject cannot honor, as CopyObject does

PutObject and CreateMultipartUpload did not check the server-side
encryption headers they were given. An x-amz-server-side-encryption
value that names no method, such as "aes:kms", matched none of the SSE
paths, so the object was stored unencrypted and the request answered
200. SSE-C together with x-amz-server-side-encryption was also accepted,
and one of the two silently won.

CopyObject already rejects both through validateEncryptionCompatibility.
Run the same check, with the same error codes, before PutObject and
CreateMultipartUpload store anything.

* s3api: answer rejected SSE headers with InvalidArgument, as S3 does

S3 rejects an unknown x-amz-server-side-encryption value and SSE-C
combined with another method with InvalidArgument. PutObject and
CreateMultipartUpload now return that code, with S3's messages, through
two new error codes; CopyObject keeps its own.

Also run ceph/s3-tests' test_put_obj_enc_conflict_c_s3, _c_kms and
_bad_enc_kms in CI, which check both handlers' responses end to end.

* s3api: close the remaining ways a PUT could skip requested encryption

- A repeated x-amz-server-side-encryption header is rejected: the
  encryption paths apply only the first value, so extra values could
  hide the method the client asked for.
- KMS options (key id, encryption context, bucket key) are rejected
  unless the method is aws:kms, matching the S3 InvalidArgument error.
- CreateMultipartUpload validates before auto-create, so a refused
  upload cannot leave a bucket behind.
- Directory markers encrypt their inline content through the shared
  SSE path and record the same entry metadata as regular objects,
  instead of storing requested-encrypted bytes in plaintext.

* s3api: read back what the SSE marker write stores

- Repeated SSE-C and KMS option headers are rejected alongside a
  repeated x-amz-server-side-encryption, closing the same first-value
  bypass for customer-key and KMS fields.
- Algorithm validity is checked before KMS options so an unsupported
  value keeps the "not supported" error.
- serveDirectoryContent decrypts marker content with the stored SSE
  metadata and returns the SSE headers, so an encrypted marker reads
  back what was written; its Content-Length now reflects the bytes
  actually served.

* s3api: HEAD of a marker skips decryption and keeps the stored size

HEAD returns no body, so it now runs only the SSE-C key check instead of
decrypting — matching HeadObjectHandler and avoiding a KMS round-trip —
and chunk-backed directory entries report Attributes.FileSize again
rather than the length of their (empty) inline content.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: keep GET Content-Length to the bytes serveDirectoryContent writes

The FileSize override described chunk-backed markers on HEAD, but GET
sends only the inline content, so it promised bytes it never wrote.

* s3api: stream a chunk-backed directory on GET like any object

A directory promoted over an uploaded object keeps the object chunks
with empty inline content, so serving only entry.Content made GET
deliver nothing while HEAD reported FileSize. GET now routes those
entries through the regular volume-server stream, keeping the two
methods consistent.

* s3api: answer a busy volume read with RequestBytesExceed

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 12:43:04 +08:00
Chris LuandDevin e3fadc6e04 filer: keep the proxy-JWT test from hanging on a failed fetch (#11664)
The handler sends the header non-blocking and the test stops on fetch
errors instead of waiting on a channel that may never be fed.

Generated with [Devin](https://devin.ai)

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 10:41:49 +08:00
Chris LuandDevin 9a9bd67efe s3api: do not re-run permission checks in the aws-chunked reader (#11655)
calculateSeedSignature called verifyV4Signature with
shouldCheckPermissions=true, which runs VerifyActionPermission — a
check that does not consult bucket policies. Every caller of
newChunkedReader is already behind the Auth middleware, which does
evaluate them, so a principal allowed only by the bucket policy was
authorized upstream and then denied when the handler built the body
reader: aws-chunked PutObject/UploadPart (botocore's default shape
over TLS, PyArrow's over HTTP as well) failed with AccessDenied.

The reader now verifies only the signature; a wrong secret is still
SignatureDoesNotMatch. The non-streaming paths already work this way.

Generated with [Devin](https://devin.ai)

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 10:01:28 +08:00
Chris LuandDevin f206d021f1 iceberg maintenance: write position-delete files with spec field ids and dictionary paths (#11652)
rewrite_position_delete_files emitted a schema-less parquet file whose
file_path column was DELTA_LENGTH_BYTE_ARRAY. PyIceberg reads delete
files with file_path as a dictionary column, which PyArrow cannot
decode from that encoding, so every table the worker rewrote failed
scans outright; readers that resolve columns by field id found none.

The row struct now declares the spec's reserved field ids
(2147483546 file_path, 2147483545 pos) and dictionary-encodes
file_path, which is also the compact choice for a column repeating one
data file's path.

Generated with [Devin](https://devin.ai)

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 10:00:48 +08:00
Chris LuandDevin 4111aa4dd8 iceberg maintenance: keep statistics files the table metadata references (#11651)
Orphan cleanup built its referenced set from snapshot manifests plus the
current and previous metadata files, then walked all of metadata/ and
data/. Table statistics (Puffin) and partition statistics files are
listed in the table metadata's statistics and partition-statistics
lists, not in any snapshot, so once they passed the safety window
remove_orphans deleted them while the metadata still pointed at them.

Mark the files in both lists referenced as well.

Generated with [Devin](https://devin.ai)

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 09:59:58 +08:00
Chris LuandDevin 9a365549c3 catalog: refused writes answer 403 to callers that can read the entry (#11650)
* s3tables: answer a denied write with 403 when the caller can read the entry

Update/Delete/Rename answered every authorization refusal as not-found so
the denial leaked no existence signal. For a caller allowed to GetTable
(or GetView) the same entry, the veil hides nothing it could not load —
yet a refused write was still answered 404, so clients saw a table they
just loaded reported as missing.

When the write check fails, re-check read permission on the same entry:
readable entries get 403 AccessDenied; invisible ones keep the
not-found answer.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg: map UpdateTable errors through writeManagerError on commit paths

CommitTable and CommitTransaction answered any non-conflict UpdateTable
failure, including AccessDenied and NoSuchTable, with a bare 500. A 500
reads as outcome-unknown to clients (PyIceberg raises
CommitStateUnknownException) where a refused commit is a plain
ForbiddenException, matching what the drop and rename handlers already
emit through the same mapper.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3tables: keep the not-found veil over views and match the real read check on renames

UpdateTable and DeleteTable read shared metadata without checking the
entry kind, so a denied write on a view reported 403 to a GetTable-only
caller where a missing name reports 404. The rename visibility check
also fed resource tags into GetView evaluation that the real GetView
path never supplies.

* s3tables: gofmt handler_table.go

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 09:59:32 +08:00
1df165d514 fix(volume): keep the TTL clock across a vacuum commit instead of rescanning (#11630)
* fix(volume): keep the TTL clock across a vacuum commit instead of rescanning

CommitCompact reloads the swapped files while holding dataFileAccessLock,
and for a vacuumed TTL volume that reload re-derived lastModifiedTsSeconds
by reading every live needle's append timestamp from the .dat: two random
reads per needle, with every read of the volume blocked behind them.

The in-memory clock is already current at that point. Every write since the
volume loaded moved it, and makeupDiff only replays writes that went through
that path. Carry it across the reload instead. This also stops an
over-budget scan from falling back to the new .dat's mtime and restarting
an expiring volume's TTL at the commit.

* volume: carry the append watermark as the TTL clock across a vacuum commit

The running append watermark is the clock the reload's recovery scan
recomputes, so the commit can keep it directly. Client-supplied needle
modified times can run ahead of or behind the append time; keeping
lastModifiedTsSeconds itself would let a forged or stale timestamp move
expiry through a vacuum, where the scan it replaces used server-side
append timestamps.

* storage: test that a vacuum commit keeps the append clock

A write's client supplied modified time can lie ahead of or behind its
append time; the commit must land the TTL clock on the append watermark,
the same value the recovery scan would have recomputed.

* volume: carry the append watermark as the TTL clock across a vacuum commit

Mirrors the Go volume server: the running append watermark is the clock
the reload's recovery scan recomputes, so the commit keeps it instead of
rescanning live needles under the write lock.

* volume: commit carries the last-write append time, not the latest append

lastAppendAtNs counts tombstone appends and is reseeded from the .dat
tail at every load, so it can sit ahead of the last write -- a delete
freshens the commit clock -- or behind it: a restarted vacuumed volume's
tail needle is not its newest write, and the commit would move the TTL
clock backward into premature expiry.

Track lastWriteAppendAtNs instead, bumped only on needle appends and
seeded by the recovery scan, so the commit lands the clock on the same
live-write maximum the rescan would have recomputed.

* volume: rescan at commit when the newest write was deleted

lastWriteAppendAtNs can hold a write the index no longer holds, so
carrying it extends the TTL clock past what recovery over the compacted
index would compute. Remember the key behind the watermark so its
tombstone or index rollback can send the reload back through
recoverLastModifiedTs, landing on the newest surviving write.

* volume: a tombstone retires the write rows beneath it in the last-write scan

A needle deleted after the compaction copy leaves its write row followed
by a tombstone in the committed index. The reverse scan skipped the
tombstone row but then counted the dead write, reseeding the watermark
and flag as if it were alive. Track keys whose latest row is a tombstone
so their earlier write rows stop counting, and exercise the
delete-inside-the-commit-window ordering in the tests.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-08 22:58:44 +08:00
Chris LuandDevin 0305e837fd iceberg maintenance: keep compacted files prunable (stats, bound order, row groups) (#11654)
* iceberg maintenance: record column statistics on compacted files

A compacted file's manifest entry was built with no column_sizes,
value_counts, null_value_counts, lower_bounds, upper_bounds or
split_offsets, so no reader could skip a compacted file on any
predicate. parquet-go already writes exact per-chunk min/max and null
counts into the footer; read that footer back after the merge and
record it on the data file, bounds as the spec's single-value
serialization with string and binary truncated as truncate(16). A
column whose bounds cannot be converted exactly gets none, and a
statistics failure is logged while the compaction commits anyway:
metrics are an optimization, not a correctness requirement.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg maintenance: merge bins in bound order so compacted files stay prunable

A bin's files were concatenated in manifest order, or largest-first
when a partition was split under the target size, so inputs disjoint
on a column came out as outputs that overlapped on it. Order each
bin's files by their bounds on one column before merging and split an
oversized partition into runs of consecutive files, so every output
covers one contiguous range. The column is the first identity field
of the table's sort order when it declares one (a descending order
sorts by upper bound), otherwise the first schema column every
candidate file has bounds for; detection resolves the same order so
it plans the bins execution builds. Files without bounds keep the old
behavior.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg maintenance: cap compacted files' row groups from table config

Neither merge writer set a row-group limit (parquet-go's default is
unlimited rows), so every compacted file was a single row group and
readers could not skip inside it either. Rows per row group now come
from the table's write.parquet.row-group-limit and
write.parquet.row-group-size-bytes, defaulting to PyIceberg's
1 048 576 rows and Iceberg's 128 MiB, with the byte size turned into
rows from the bin's inputs' compressed bytes per row and a floor of
1 024 rows. Both writers take the cap, and the statistics the entry
records list one split offset per row group.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg: keep compaction order eligibility per group, rescue stranded runs

The merge order resolved over all candidates, so one oversized or
non-Parquet file without bounds disabled ordering for files that could
participate. Resolve it per partition group over the eligible entries.

Ordered runs too short to merge were dropped entirely. Runs from an
inferred bounds order now fall back to size-based packing — ordering is
a preference there — while runs under a declared sort order are still
left for later passes so the sort contract holds.

An explicit write.parquet.row-group-limit is a cap, not a floor: values
below the estimate floor are now honored instead of being raised.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg: only use the declared sort order when its bounds are complete

An entry without bounds on the sort column sorted to the tail and merged
into an output claiming an order it cannot verify. Fall back to
bound inference instead of ordering by a later sort field alone.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg maintenance: keep ordered runs when the full repack yields nothing

The bestEffort leftover fallback removed the ordered runs before checking
whether repacking the whole bin produced any bins, discarding valid
compaction work.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 22:06:44 +08:00
Chris LuandDevin 8f80dac30f ec: strict_placement option so encode only runs while guarantees hold (#11656)
* ec: strict_placement option so encode only runs while guarantees hold

Shard placement during encode was best-effort (PlaceDurabilityFirst):
when the cluster could not satisfy the per-disk caps, anti-affinity,
replica-placement or per-rack caps, the constraints were relaxed and
the volume was encoded anyway, weaker than configured. A
strict_placement option on the erasure coding task switches planning
to PlaceStrict so the volume's planning fails instead, and the encode
is retried when capacity allows the guarantee.

Also documents the resilience rule in ec.encode help: a volume
survives losing any nodes or racks holding at most parity-shards
shards between them, and how -shardReplicaPlacement's rack and node
digits bound that loss.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ec: expose strict_placement through the plugin form and persisted task policy

The admin UI, the admin.toml maintenance mapping, and the TaskPolicy
serialization all dropped the new flag; add the bool field to
ErasureCodingTaskConfig, the worker config form, and both conversion
directions.

* shell: describe shardReplicaPlacement as requested limits, not guarantees

ec.encode places shards best-effort, so the configured rack/node caps only
bound shard loss when the final placement actually satisfies them.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 22:05:02 +08:00
Chris LuandDevin 4d1f49c638 filer.sync: sign proxied chunk I/O from the per-side security file (#11645)
* filer.sync: sign proxied chunk I/O from the per-side security file

The -a.security / -b.security files were used for gRPC TLS and the
HTTPS client but not for jwt.filer_signing, so filer-proxied chunk
reads and writes carried a token signed with the process-wide key and
failed authorization whenever the two clusters' keys differ.

LoadFilerJwtFromFile returns a FilerJwtProvider for each side's file,
which FilerSource and FilerSink now accept for proxied chunk reads and
writes. With no keys in the file or no flag, both fall back to the
process-wide jwt.filer_signing configuration as before.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* replication: use the side filer read key for manifest downloads and fall back per access level

Manifest chunk resolution still signed proxied downloads with the
process-wide read key, so a source filer requiring its own key 401'd on
manifest-bearing files. A side security file that set only one access
level also produced empty tokens for the other instead of inheriting the
process-wide key, and the side file loader ignored the WEED_ environment
overrides the filer itself honors.

ResolveChunkManifest/ResolveOneChunkManifest keep their signatures;
FilerJwt-aware variants thread the provider down to fetchWholeChunk,
which prefers it on proxy URLs. The side loader now applies the same
environment precedence and falls back to the process-wide signer per
missing access level.

* security: verify the configured filer token lifetimes

* security: reject negative filer token lifetimes

A negative expires_after_seconds reached GenJwtForFilerServer and produced
a token with no expiration claim. Also synchronize the Authorization-header
capture in the proxy test and restore the prior viper key on cleanup.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 22:04:06 +08:00
Chris LuandDevin 14fdd61aea filer: stop the aggregated metadata subscribe loop rescanning an exhausted persisted log (#11644)
* fix(filer): gate the aggregated metadata disk pass on real change

A subscriber whose start position is past the end of the local persisted
log re-ran the whole persisted-log pass - store listings, file opens,
readahead - on every loop iteration. Each iteration is paced only by the
shortest wake (the 20ms hold floor on a busy watermark), so one parked
subscriber kept a full CPU core busy for the life of the stream.

The aggregated loop now mirrors the local loop's gate: the disk pass
runs on the first pass and afterwards only when something it cannot
miss changed - a local flush landed, the peers' flush low-watermark
advanced (more content admitted, or new files in a shared store), the
cursor moved, or a disk hold is pending (the ring read that follows an
empty pass parks internally, so skipping there would strand a held
entry).

Regression test: a subscriber parked past the persisted-log tail holds
the listing rate near zero and still delivers once peers report
progress.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: re-arm the aggregated disk pass on unobserved change

Review found three staleness classes the gate could not see: the flush
low-watermark only catching rises (a joining peer lowers the minimum and
invalidates an earlier pass's proof), a peer past the minimum landing a
file without moving it, and a chunk subscriber's refs-stop bound
advancing with wall time. Re-read when the low-watermark moves in either
direction, when the chunk listing bound admits more files, and on a slow
re-probe cadence for files no watermark can signal.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: unwind the parked ring read so the disk re-probe runs, and re-read on cursor rewinds

A caught-up subscriber parks inside LoopProcessLogData's wait loop, so
the re-probe interval in the outer disk gate could never elapse there;
the callback now unwinds the read once the cadence is due so the gate
re-evaluates. The cursor trigger also needs to notice rewinds, not just
advances, since ResumeFromDiskError moves the cursor backward.

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 22:03:31 +08:00
Chris LuandDevin 56fbc3deb5 s3api: align S3/IAM error responses with AWS (#11632)
* s3err: add InvalidArgument and AuthorizationHeaderMalformed codes

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: answer unrecognized bucket PUT sub-resources with 501

A PUT on a bucket carrying an unrecognized query (logging, metrics,
intelligent-tiering, ...) fell through to the bare CreateBucket route
and returned BucketAlreadyOwnedByYou or re-created the bucket. AWS
answers these with NotImplemented.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: reject malformed copy-source and multipart PUT parameters

A malformed X-Amz-Copy-Source or a non-numeric partNumber fell through
to the plain PutObject route and stored the body as a regular object.
Answer them with InvalidArgument-class errors instead of writing data.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: verify x-amz-content-sha256 against the streamed body

A PUT carrying a hex or base64 payload hash now streams through a
verifier that reports a mismatch once the stream is exhausted, instead
of storing an object that does not match its declared hash. The error
is deferred so intermediate reads that drop (n>0, err) results cannot
silently swallow it. Sentinel values (unsigned/streaming payloads)
remain exempt and malformed values fail fast.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: map truncated SigV4 headers to the error for the missing field

AWS answers an Authorization header missing Credential= with
InvalidArgument and one missing or malformed Signature= with
AuthorizationHeaderMalformed, instead of a generic MissingFields.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: answer IAM/STS failures in the query-protocol envelope

Embedded IAM and STS routes now report authentication, form-parse and
authorization failures with the IAM ErrorResponse body instead of the
S3 Error envelope, so IAM SDK clients can parse them. Requests signed
for s3 keep the S3 envelope, keyed off the credential scope.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: return 403 AccessDenied when the request has no Date

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: answer throttling rejections as SlowDown

ErrTooManyRequest and ErrRequestBytesExceed reported made-up codes;
AWS serves these throttling rejections as SlowDown.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: reject versionId requests on buckets that never had versioning

GET, HEAD and DELETE carrying a non-empty versionId on an unversioned
bucket now fail with InvalidArgument instead of being answered as a
plain object request.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: answer DeleteObjects over 1000 keys with MalformedXML

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: InvalidArgument for non-numeric or out-of-range part numbers

partNumber=abc, 0 and >10000 all resolve to InvalidArgument, matching
AWS, instead of InvalidPart.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: correct LimitExceeded, InvalidAction and ServiceFailure mappings

LimitExceeded is a conflict (409), an unknown Action is InvalidAction
(404) rather than NotImplemented, and internal failures report the IAM
receiver fault type. Applies to both the embedded IAM endpoint and the
standalone iamapi server.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: refuse DeleteUser while access keys remain

Deleting a user with live credentials orphaned its access keys; AWS
answers DeleteConflict until they are removed first.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test/s3: add S3/IAM error-response compatibility harness

* s3: return InvalidArgument for malformed x-amz-content-sha256

A header value that decodes to neither 32-byte hex nor base64 is a
malformed argument, not a hash mismatch.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: stop spawned mini when readiness times out

A slow-starting server otherwise survives the failure path and keeps
the S3 port occupied for the next run.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: register atexit cleanup before setup

A failed setup previously skipped cleanup, leaking the bucket and IAM
user on persistent servers.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3api: reject unrouted subresources on the DELETE bucket catch-all

PutBucketHandler gained the same guard when the route-level check moved
into the handlers; DeleteBucketHandler was missed, so an authorized
DELETE /bucket?logging could delete the bucket instead of answering
NotImplemented.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: tolerate unset fixture variables in cleanup and cover DELETE ?logging

Cleanup now runs its IAM/multipart steps only when setup reached them, so
an early setup failure still removes the bucket. Added a DELETE
bucket-subresource case asserting NotImplemented and bucket survival.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: fail a case when its side-effect check reports a regression

A non-empty check note now fails the case, so a deleted bucket or an
object created by a malformed request cannot slip through behind a
passing status check.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 21:02:38 +08:00
Joel Town RoadandChris Lu 079a7d8ba3 s3: give each prefix its own hidden-prefix probe budget (#11626)
* s3: give each prefix its own hidden-prefix probe budget

The shared cursor.probedEntries counter in dirHoldsOnlyHiddenEntries was
exhausted by one large all-deleted subtree, causing every later prefix in
the same request to be treated as visible.  The fix allocates a fresh
budget (hiddenProbePerPrefixBudget, default 1000) for each top-level call
and threads it down to recursive calls via a pointer, so cross-prefix
budget bleed is impossible.

Fixes #10847.

* s3: keep 10000 probe budget, now per prefix

Restore hiddenProbeBudget = 10000 as a package const (not a mutable var)
applied per-prefix instead of per-request. The override hook moves to an
unexported probeBudget field on ListingCursor; zero means use the package
default. Tests set probeBudget: 3 on the cursor, keeping them fast without
touching package state. Also revert the unrelated uint32 cast on the
ListEntries Limit field.

* s3: cap total hidden-prefix probe work per listing request

The per-prefix budget resets for every candidate prefix, but deleted
prefixes do not spend maxKeys, so a page can walk an unbounded number of
them. Keep a request-wide probedEntries ceiling (hiddenProbeTotalBudget,
10x the per-prefix budget) so total probe work stays bounded.

* s3api: pin the probe-budget cutoff and the request-wide ceiling

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-07 20:53:39 +08:00
Eliah RusinandClaude Opus 5.5 0bcebf708c ecbalancer: cap total shards per rack in Plan (#11623)
* ecbalancer: cap total shards per rack in Plan

Plan caps data and parity per rack separately (ceil(data/racks) and
ceil(parity/racks)), so with 10+4 over 8 racks a rack can legally hold
2 data + 1 parity. When a rack is one disk, losing two such racks loses
6 of 14 shards and the volume can't be read.

The cross-rack phase now also caps each rack's TOTAL shards of a
volume, sized with Place's rackTotalCap: ceil(shards/racks) unless the
racks lack room, counting a rack's own shards of the volume as room
since Plan can move them.

- A rack above the cap sheds parity until it fits. Those shards may go
  to a data-bearing rack, and when no rack is under the parity cap
  they fall back to a rack under the total cap. #11438's non-overflow
  candidates keep moving only to data-free racks.
- No cross-rack move lands on a rack at the cap.
- A rack above the cap triggers balancing regardless of the imbalance
  threshold.
- The fallback applies only while the source rack is above the cap;
  otherwise the next Plan moves the shard back.
- Options.RackTotalCapRaised reports volumes whose cap had to be raised
  above the even share; the worker and the shell log it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ecbalancer: finish the rack cap in one Plan, count SameRackCount room

Review follow-ups.

- The data pass can run out of destinations before the parity pass frees
  slots elsewhere, which left a data-heavy rack above the cap after one
  Plan (a one-shot shell balance stops there). The cross-rack phase now
  repeats while a rack is above the cap and the last round moved
  something. A shard moves at most once per plan, since each move runs
  as its own task.
- planRackTotalCap bounds each node's room by what SameRackCount still
  allows, as Place does. Counting the raw free slots sized the cap too
  low, so RackTotalCapRaised missed volumes it should report.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 20:35:25 +08:00
576837f1b2 fix(s3): make self-heal pointer persist CAS-bound against concurrent writers (#11627)
* fix(s3): make self-heal pointer persist CAS-bound against concurrent writers

Follow-up to #11618: pointerless reads of a slash key whose regular-path
entry is a physical parent (or a bare-key object) now fall through to
healStaleLatestVersionPointer, which rescans .versions and persists a
repaired pointer. The persist was an unconditional upsert off the pre-scan
snapshot, so a PUT or delete that atomically advanced the pointer on the
owner filer while the heal was rescanning could be rolled back, making
older content or ACLs current again.

Mirror the CAS discipline clearStaleLatestVersionPointer already applies:
re-fetch the live .versions entry, require its pointer fields to still
match the ones the heal observed, and abandon the persist (still returning
the rescanned entry) when a concurrent writer has moved them. Write the
live Extended map so concurrently updated fields are preserved.

* fix(s3): close the check-then-act window in the self-heal pointer persist

The CAS re-fetch added in the previous commit narrows the race but leaves
a gateway-side window: after the live .versions entry is re-read and the
pointer compared, the repair is still written back through an
unconditional RPC, so a PUT or delete committing between the re-fetch and
the persist still ends up rolled back by the stale repair.

Bind the persist to the live image the heal just re-read with an
IF_ENTRY_EQUAL precondition, the same discipline routedSelfCopy applies
to stale self-copies: the filer evaluates the condition under the entry's
path lock and conditional writes route to the owner filer, so a writer
committing inside the window fails the precondition and the winner's
pointer stands. FailedPrecondition and NotFound are authoritative replies
and are not replayed by the failover layer.

The test now also covers a writer committing during the persist, which
reverts the pointer on the previous unconditional write-back.

* s3api: CAS-bind the stale-pointer clear against concurrent writers

The pointer clear re-read the live .versions entry and then wrote it
back unconditionally through mkFile, so a writer committing between the
re-fetch and the persist was rolled back to a cleared pointer. Persist
through the same IF_ENTRY_EQUAL conditional update as the repair path.

* s3api: test the CAS contract on the stale-pointer clear

* s3api: never clear a pointer the clear did not observe as stale

The CAS clear skipped its live-pointer match when the caller's snapshot
carried an empty latest-version id, so a writer promoting a version
between the clear's rescan and its re-fetch had the fresh pointer
CAS-cleared away (expected = the writer's own live entry), briefly
making the just-written version appear absent. With an empty observed
id, reaching the persist at all implies a concurrent promotion (an idle
key short-circuits as already-clear), so make the pointer match
unconditional and abort instead.

Extend TestClearStaleLatestVersionPointerConcurrentWriter with
pointerless-snapshot cases: a post-rescan promotion must survive, and an
idle pointerless key must short-circuit as already-clear. The fake
filer's proto round-trip drops empty Extended maps, so the snapshot is
padded the way real callers do.

---------

Co-authored-by: zhaoyuchen <yc.zhao@yinzon.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-07 20:04:09 +08:00
Chris Lu ec261c5fbc S3: quote ETag in CopyObject and UploadPartCopy XML responses (#11624)
* s3api: extract quoteETag from setEtag

Consolidate ETag quoting so the XML response builders can share it.

* s3api: quote ETag in CopyObjectResult XML

AWS returns the ETag quoted in the copy result body, matching the ETag header. Fixes seaweedfs/seaweedfs#11622

* s3api: quote ETag in CopyPartResult XML

UploadPartCopy returned the raw ETag in the XML body while the response header and other APIs return it quoted. Fixes seaweedfs/seaweedfs#11622

* s3api: test ETag quoting in copy responses

* s3api: assert quoted ETag on the wire in copy response tests

* s3api: use strconv.Quote in quoteETag

Addresses CodeQL 'potentially unsafe quoting' on string concatenation.
2026-10-06 18:05:15 +08:00
zhao-ycandChris Lu d829275de6 fix(s3): recheck cache metadata and validate null objects (#11618)
* fix(s3): recheck cache metadata and read null directory markers

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* fix(s3): preserve null object identity and bound test cache

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* s3: trim comments on the anonymous read cache path

Signed-off-by: Chris Lu <chris.lu@gmail.com>

---------

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>
Signed-off-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-06 14:25:43 +08:00
Joel Town RoadandChris Lu 23893eb378 volume: return error instead of panicking when .dat open fails (#11619)
* volume: return error instead of panicking when .dat open fails

When backend.OpenVolumeFile returns an error (e.g. disk below
-minFreeSpace), dataFile is nil.  Calling backend.NewDiskFile(nil)
immediately after caused a nil-pointer panic in f.Stat()/f.Name().

Move the existing error check to run right after OpenVolumeFile, before
NewDiskFile is called, so the error is returned cleanly.

Fixes seaweedfs/seaweedfs#11615

* volume: share .dat load error handling

Extract datFileLoadError helper so open and create paths share one check.

* volume: trim dat-open-fail test comments

* volume(rust): cover unopenable .dat load path

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-06 14:24:42 +08:00
32e77ff980 Secure Weed Mini Admin Listeners by Default (#11613)
* securing admin

Signed-off-by: Subhadeep Maity <smaity@slb.com>

* updated readme

Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>

* fixed pr comments

Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>

* docs: tidy weed mini admin bind notes

Drop the new single-entry CHANGELOG.md since changes are documented via
GitHub releases, and rewrap the README paragraph to match the surrounding
one-line style without self-referential issue/PR links.

* review comments

Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>

---------

Signed-off-by: Subhadeep Maity <smaity@slb.com>
Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>
Co-authored-by: Subhadeep Maity <smaity@slb.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-06 13:33:47 +08:00
dependabot[bot]andChris Lu de75450655 build(deps): bump github.com/apache/iceberg-go from 0.6.1-0.20260817192109-c2105090c9e2 to 0.7.0 (#11609)
* build(deps): bump github.com/apache/iceberg-go

Bumps [github.com/apache/iceberg-go](https://github.com/apache/iceberg-go) from 0.6.1-0.20260817192109-c2105090c9e2 to 0.7.0.
- [Release notes](https://github.com/apache/iceberg-go/releases)
- [Commits](https://github.com/apache/iceberg-go/commits/v0.7.0)

---
updated-dependencies:
- dependency-name: github.com/apache/iceberg-go
  dependency-version: 0.7.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* iceberg: write delete manifests with NewManifestWriter

iceberg-go 0.7.0 validates that delete entries only land in
delete-content manifests, so the WriteManifest + byte-level content
patch workaround no longer works: WriteManifest always creates a
data-content writer and rejects delete entries outright.

Use NewManifestWriter with WithManifestWriterContent instead and
dispatch entries through Add/Existing/Delete by status. The returned
ManifestFile now carries writer-computed counts, partitions and
min-sequence-number, so the manual ManifestFile rebuild is dropped.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-06 13:32:54 +08:00
c739e5cf78 fix(s3): drain in-flight requests on SIGTERM in standalone weed s3 (#11603)
Standalone weed s3 only shut its servers down when shutdownCtx was set,
which only weed mini does. On SIGTERM the interrupt hooks ran and the
process exited with requests still in flight, so clients saw connection
resets during rolling restarts.

Register an interrupt hook that drains every S3 HTTP(S) listener
(including the local and unix socket ones, and the Iceberg and Lance
servers) for up to 15s alongside a bounded gRPC GracefulStop, then closes
the S3 API server, reusing the filer's shutdown helper. Serve exits now
join that shutdown so the process does not exit mid-drain, and the
secondary listeners tolerate http.ErrServerClosed instead of exiting
fatally.

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:10:37 +08:00
Peter Dodd 52c6df3bec fix(filer): leave the lock ring before stopping gRPC on shutdown (#11604)
A draining filer stayed in the master's lock ring until its process
exited, while gRPC GracefulStop was already refusing new connections. S3
gateways and peer filers kept routing object-write locks and owner-routed
writes to it for the whole graceful-stop window.

On shutdown the filer now sends leave_lock_ring on its open KeepConnected
stream. The master removes it from the lock ring only; it stays a cluster
member so peers keep following its metadata log through the drain. Once
the ring update without it arrives, the filer has already transferred its
locks to the new owners, and it keeps serving through the prior-owner
window (plus a second for peers that apply the update later) before gRPC
and HTTP begin draining. The whole leave is bounded at 10s so slow lock
transfers or a stuck stream cannot hold up the drain. A lone filer, or a
master that ignores the message, falls back to the previous behavior.
2026-10-06 09:53:40 +08:00
aa5b337716 fix(s3): honor object ACLs for anonymous GET and HEAD (#11605)
* fix(s3): honor object ACLs for anonymous reads

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* s3: drop unsigned session tokens on deferred anonymous reads

An unsigned GET/HEAD carrying X-Amz-Security-Token is anonymous, not a
session request; strip the token before authentication and authorization
so it cannot influence identity or policy evaluation.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:52:37 +08:00
Chris LuandDevin 076fd24186 filer: bound metadata log flush retries during shutdown (#11596)
* filer: bound metadata log flush retries during shutdown

On SIGTERM the filer could hang forever in Shutdown: the final
LocalMetaLogBuffer flush retries appendToFile indefinitely, and with
the master already down each AssignVolume attempt just kept failing.
WaitForShutdown never returned, the interrupt hook never reached
os.Exit, and the half-dead filer kept its ports bound.

Thread a context through appendToFile/assignAndUpload and switch to
a 15s-bounded context once the filer is stopping: the flush abandons
with a log line instead of retrying forever. Normal operation keeps
the unbounded retry so no metadata is dropped while running.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: share one shutdown deadline across all pending meta log flushes

Review feedback on the per-flush timeout: a deadline armed at flush start
could already be expired when shutdown arrived, the first append of a flush
still ran unbounded, each queued window got a fresh budget (16 windows *
15s), and an abandoned window still advanced the flushed watermark as if it
had landed.

Rework to a single shared flush context on the Filer, cancelled once by
Shutdown via AfterFunc. Every append - the in-flight one and every queued
window - observes the same deadline, so the whole drain is bounded at 15s.
A flush that gives up reports its dropped bytes through the new
LogBuffer.NoteFlushDropped, and loopFlush then skips the offset/timestamp
advance and subscriber notifications for that window.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound append attempts and commit uploaded pieces detached

Store operations check then drop request cancellation, so a stalled
backend could still hold flushFn past the shutdown deadline; run each
append attempt on its own goroutine and give up on it at the deadline.
Once a piece is uploaded, commit its entry on a detached context so the
expired deadline cannot strand the chunk.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: keep shutdown flushes synchronous

Detaching the append attempt let flushFn return while the goroutine
still held the pooled flush buffer and could commit after the metadata
store closed; abandonment is only safe for the cancelable assign/upload
phase, which the shared flush context already bounds.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:51:02 +08:00
a7a590474d kafka: honor notification.kafka.event_types (#11601)
* kafka: honor notification.kafka.event_types

Kafka published every filer event and ignored the filter the webhook notifier already uses.

Co-authored-by: Cursor <cursoragent@cursor.com>

* notification: share event-type classification between queues

Kafka duplicated the webhook's event classification verbatim; move it to
the notification package so the two queues cannot drift. Webhook keeps
its typed eventType wrappers over the shared helpers.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:15:06 +08:00
8c67b75190 image: add an optional public image processing gateway (#11593)
* image: add an optional public image processing gateway

* image: fix representation metadata and processing bounds

* image: restrict passthrough to non-executable media types

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* image: tighten source media-type validation

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* image: write passthrough body on the inner response writer

CodeQL still flagged the passthrough write: the content type was set on
the wrapper while the body reached w.ResponseWriter, so the validated
header could not be associated with the write. Set headers and copy the
body on the same inner writer.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* image: serve processed output on the inner response writer

* image: reject XML source types and unsafe conditional metadata

---------

Co-authored-by: zhaoyuchen <yc.zhao@yinzon.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 19:49:02 +08:00
zhao-ycandChris Lu 16e66b1bad fix(s3): initialize destination ACLs for CopyObject (#11599)
* fix(s3): initialize destination ACLs for CopyObject

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* fix(s3): re-check routed self-copy eligibility on the locked read

routeInPlace was decided on the pre-lock entry, but the PATCH body
re-reads the entry. A concurrent write changing file mode or MIME in
between left the routed PATCH installing new ACL keys while Attributes
kept the stale mode. Evaluate eligibility against the re-read entry and
retry the self-copy under the distributed lock when it no longer
qualifies.

* fix(s3): guard metadata self-copies against concurrent writes

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* ci: raise s3api unit-test timeout to 9m

The suite crossed the 5m binary timeout on the hosted runner (local run
is ~4.3m and still growing). The job-level limit is already 10m.

* ci: allow setup time before the S3 API test suite

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

---------

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-05 18:22:16 +08:00
Chris Lu 7e809c9991 admin: stop leaking cancelled maintenance task files in -dataDir/tasks (#11597)
* admin: delete persisted state when scan cancels pending tasks

Each detection cycle cancels every pending task of a type before
re-detecting it, and the cancel path saved the cancelled task back to
disk. Nothing ever removed those files, so -dataDir/tasks gained one
orphaned .pb per candidate volume per scan cycle.

Cancelled is terminal, so drop the file the same way CompleteTask does
for completed/failed tasks. The cancelled entry stays in memory for the
UI until the next purge.

Refs #11595

* admin: delete persisted state when CancelTask cancels a pending task

The manual cancel path only updated memory, leaving the pending .pb on
disk where a restart would resurrect the cancelled task as pending and
the file would linger until then. Delete it like the scan-cycle cancel
path now does.

* admin: count cancelled tasks toward task retention cleanup

CleanupOldTasks and ConfigPersistence.CleanupCompletedTasks only
filtered completed/failed tasks, so cancelled entries were exempt from
retention in both memory and on disk. Treat all terminal states alike;
nil CompletedAt entries also count and sort last, so they are pruned
first.

* admin: run task file retention in the periodic cleanup loop

cleanupCompletedTasks had no callers, so the on-disk retention bound
never ran during uptime. Invoke it from performCleanup alongside the
in-memory CleanupOldTasks sweep.

* admin: guard task state writes against stale saves and failed deletes

saveTaskState runs after mq.mutex is released, so the task may have gone
terminal in between; a delayed pending save could then recreate the file
a cancel just deleted and resurrect the task on restart. Skip saving
non-terminal snapshots once the live task is terminal or gone.

If a cancel file removal fails, fall back to writing the cancelled
snapshot so the file is terminal rather than pending. deleteTaskState now
returns its error, and CancelTask captures task.Status while still
holding the queue lock.

* admin: serialize task file check+write against cancel deletes

The saveTaskState guard still had a check-then-write window: a pending
snapshot could pass the terminal check before a cancel deleted the file,
then write it back after. A persistMu on the queue now covers the
check+save and the cancel paths' delete (with its terminal-state
fallback), so the two cannot interleave for the same task.
2026-10-05 12:02:53 +08:00
Chris Lu 3c17c5146e S3: fix ListObjectVersions losing keys across page boundaries (#11598)
* s3api: thread filer client through the versioned-listing collector

findVersionsRecursively now binds one SeaweedFilerClient for the whole
recursive walk instead of re-resolving a filer on every list/lookup call,
and the collector's list/getEntry/scanLatestVersionEntry/getObjectVersionList
helpers go through it. No behavior change; this also lets tests drive
collectVersions with a stubbed client.

* s3api: keep collecting versions while pending names can sort into the page

ListObjectVersions walked the filer in directory-entry name order and
stopped as soon as maxKeys+1 items were collected, sorting only that
partial set. Filer names do not match key order: "a.copy.versions" sorts
before "a.versions" while key "a.copy" sorts after "a", so a page
boundary inside the earlier-walked sibling's versions permanently skipped
the later key.

Track the largest key collected (maxKey) and, once the collector is full,
keep walking until entry names pass the ceiling of names that can still
resolve to keys at or below it; the ceiling reaches through the prefix
versions of maxKey. Entries whose subtree can only hold keys above maxKey
are skipped. Versions of an in-bound object are collected in full so its
position in the sorted page is exact.

Fixes seaweedfs#11594

* s3api: resume versioned listings at the earliest covering name prefix

computeStartFrom mapped the key marker straight to an entry name (or cut
it at the first '/'), which skips sibling directories that are a prefix
of the marker below '0' - for marker "d.x" the listing resumed at name
"d.x", skipping directory "d" whose keys "d/*" all sort after it.

Resume at the earliest remainder prefix ending at a byte below '0' ('/',
'.', '-' and friends), so every directory whose subtree can still hold
keys past the marker is revisited; already-returned keys inside are
filtered by the existing marker checks as before.

* s3api: regression test for versioned-listing pagination order

Drive collectVersions against a stubbed filer holding the issue-11594
layout - "a.copy.versions" listing before "a.versions", plus a "d/"
subtree next to "d.x" - and assert that every page size from 1 up
reproduces the unpaginated ordering with no lost or duplicated entries.
Also updates TestComputeStartFrom for the new earliest-prefix resume and
gives testFilerClient a LookupDirectoryEntry stub.

* s3api: inject list/getEntry functions into the version collector

Pinning one SeaweedFilerClient for the whole walk dropped per-call
failover: previously each s3a.list resolved a filer through
WithFilerClient, so a mid-walk filer failure could fall back to a
healthy peer. Inject s3a.list/s3a.getEntry as function fields instead -
production keeps the failover behavior, tests can still stub.

* s3api: keep scanning marker for a later covering prefix

A leading byte below '0' (marker .hidden/file) has no non-empty prefix
at index 0, but a deeper separator still does - resuming at .hidden/file
skipped the .hidden directory and its keys after file. Continue the scan
instead of bailing on the first byte.
2026-10-05 11:58:04 +08:00
zhao-ycandChris Lu 483dd4b12e s3api: persist ACLs on PutObject uploads (#11592)
* s3api: persist ACLs on PutObject uploads

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* s3api: fix PutObject ACL edge cases found in review

- Only enforce BucketOwnerEnforced when explicitly configured; buckets
  without a stored ownership control keep accepting upload ACLs
- Ignore ACL query parameters on SigV2 requests, which do not sign them
- Mirror signed-query ACL values into headers after authentication so
  grant parsing and resolveFileMode agree on presigned uploads
- Validate only caller-supplied grantees against the account registry;
  default grants now work for accounts outside the local registry
- Reject unknown grantee keys and accept comma-separated grantee lists
  without spaces in ParseCustomAclHeader
- Guard against identities without an account

* s3api: harden upload ACL parsing and authorization

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* s3api: evaluate upload ACL grantees individually in policies

A comma-joined grant header or a signed query parameter reached policy
conditions as one value, so a deny on a later grantee did not fire. Split
grant headers into per-grantee values for policy evaluation and share the
grantee pair parser with ParseCustomAclHeader.

* s3api: keep raw grant header values visible to policy conditions

Exact-match conditions written against the signed header value stopped
matching once grantees were split for evaluation. Preserve the original
wire values alongside the per-grantee values so deny policies fire on
either granularity.

* s3api: evaluate upload ACL grants as one canonical list in policies

Conditions on s3:x-amz-grant-* now see a single comma-separated canonical
grant list identical for a single line, repeated header lines, or a signed
query parameter. This keeps StringEquals allows and exact-list or
allowlist (StringNotEquals) denies accurate regardless of wire encoding.

* s3api: preserve upload ACL denies and align policy checks

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* s3api: retain upload owner grants and literal policy values

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

---------

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-05 09:13:19 +08:00
9d1c24d80d [Filer] Support append to inline small files (#11591)
* fix 11586

* Update filer_server_handlers_write_autochunk.go

* filer: fix inline append races, empty files, and stale ETags

Serialize the append read-modify-write on the entry lock so concurrent
appends merge instead of losing content, keep small appends to empty
files inline, tolerate legacy entries whose metadata size differs from
their content, and set the entry digest so appended inline files keep a
real ETag.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 09:12:20 +08:00
Chris Lu d9b69a7f76 s3api: fix PutObjectAcl permission scoping and owner grants (#11587)
PutObjectAcl had four authorization and ownership bugs:

- The handler embedded the resource path into the action
  (WriteAcp:bucket/object), and authRequest/CanDo then scoped it to
  the request's bucket/object again. A bucket-wide WriteAcp:bucket
  grant could never match, so legitimate owners got 403.
- After authRequest succeeded via an IAM or bucket policy, a leftover
  identity.CanDo gate re-checked only the legacy Actions list, denying
  identities authorized purely by policies.
- For canned and default ACLs, ExtractAcl generated the FULL_CONTROL
  grant for the requesting account instead of the object owner. An
  admin setting private/public-read on another account's object left
  the owner metadata intact but reassigned full control to the admin.
- Objects without stored owner metadata (e.g. written via the filer
  outside S3) fell back to treating the requester as the owner, so any
  user with a WriteAcp grant could take them over. Non-admins are now
  denied; admins keep the takeover fallback.

Grantee validation now also accepts the object's stored owner even when
that account has been removed from the registry, so canned/XML ACLs for
retired owners keep working.
2026-10-04 19:41:10 +08:00
Chris Lu 2b5fdc639f filer: stop isSameChunks from sorting caller-owned chunk slices (#11584)
* filer: stop isSameChunks from sorting caller-owned chunk slices

slices.SortFunc reorders the input in place. filer.remote.sync calls IsSameData on a metadata event's NewEntry inside isMetadataOnlyUpdate and later stamps the filer entry under an IF_ENTRY_EQUAL precondition carrying that same entry. The ETag-sorted chunk list never matches the stored entry, so every stamp of a multi-chunk object fails, synced_mtime_ns stays zero, and dirty objects are re-uploaded forever. Sort clones of the slices instead.

* filer: test IsSameData leaves input chunk order unchanged

Guards the clone-then-sort fix: a regression back to in-place sorting would reorder caller-owned chunk slices and reintroduce the remote-sync IF_ENTRY_EQUAL mismatch.
2026-10-04 17:37:26 +08:00
39bc9cd0ef s3api: copy the trailer checksum before reading the next trailer line (#11583)
* s3api: copy the trailer checksum before reading the next trailer line

parseChunkChecksum kept the checksum value as a sub-slice of the line
returned by bufio.Reader.ReadSlice, which is only valid until the next
read. When the trailer lines arrive in separate TCP segments, reading
x-amz-trailer-signature refills the buffer and overwrites the saved
value, so a correct upload fails with InvalidDigest ("The Content-Md5
you specified is not valid").

The AWS SDK for Java v2 (>= 2.30) on a Linux JDK sends the trailer that
way; about half of its signed streaming uploads failed.

Fixes #11582

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* s3api: reuse crc32 writer and trim comments in trailer split test

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-04 01:14:15 +08:00
Chris Lu f4ef37e752 filer.sync: resubscribe the metadata stream when a failure pins the offset (#11581)
* filer sink: keep the gRPC status inside wrapped errors

%v stringifies the status, so a peer teardown reported as Canceled ("the
client connection is closing") reached IsTransientError as plain text and
matched nothing: the sync job failed on the first attempt and pinned the
offset. %w keeps the status reachable, so the retry runs on a fresh
connection once the target is back.

* pb: let a consumer drop the metadata stream to force a resubscribe

A MetadataProcessor job that exhausts its retries pins the processed
watermark so the event replays on the next subscribe — but nothing on the
source stream notices a target-side failure, so the replay waited for an
unrelated reconnect or a restart. The new Resubscribe channel cancels the
stream's context; the Recv loop answers it with ErrResubscribe so the
caller's retry loop resubscribes from GetResumeTsNs and replays the pinned
events in order.

* pb: stop the event retry loop once the stream context is done

RetryUntil ignores context, so a subscriber parked on a failing offset
write would keep retrying past a resubscribe signal until the sink came
back. Stop retrying when the stream is being dropped so the resubscribe
takes effect promptly.

* filer.sync: signal resubscribe when a job failure pins the offset

A job that exhausts its in-job retries leaves the event pinned behind oldestFailedTsNs, replayable only on a reconnect. Closing resubscribeCh on the first recorded failure lets the metadata follower drop the stream so the reconnect replays the pinned events instead of waiting for a process restart (#11572).

* filer.sync: wire the resubscribe signal into the follow options

filer.sync, filer.remote.sync, and the remote gateway bucket sync all run their subscription inside an outer retry loop, so ErrResubscribe resurfaces as a resubscribe from the persisted watermark.

* filer.sync: wait for in-flight jobs before signaling resubscribe

* remote sync: never resume past the saved offset when -timeAgo is set

* filer.sync: drop events that arrive after the drain signals resubscribe

* pb: interrupt the event retry backoff when the stream context ends

* filer.sync: stop admitting once a failure pins, and count jobs per timestamp

A pinned watermark only released once the processor went fully quiet, so a busy stream could starve the resubscribe — the failed event would wait for an unrelated reconnect anyway, the wait this mechanism exists to remove. The processor now latches stopped when a job fails: admission drops new events (they replay from the pinned watermark after the reconnect), a broadcast releases blocked waiters, and the resubscribe signals as soon as the jobs already in flight drain. A redelivery of an event still in the failure ledger may still run so its success shrinks the replay, but nothing starts once the signal has fired, or it would race the replay it asked for.

Dropped events no longer inflate the received counters — an event counts only once admitted, and the replay's own admission counts it.

While here: activeJobs keyed by TsNs collapsed events sharing a timestamp, so one completion could empty the map while a same-ts sibling was still running — letting the drain gate and the watermark outrun it. Jobs are now counted per timestamp, and the drain and lazy heap cleanup go through the counts.
2026-10-04 00:06:55 +08:00
10b0f2b8ad volume server: refuse the rest of a grouped run after a durable index failure (#11576)
* volume server: refuse the rest of a grouped run after a durable index failure

A durable write whose needle-map put fails stops the volume taking writes
(#10825): sent on its own, the next write then fails read only before it
appends. The grouped run from #11543 appends and syncs every entry before
publishing any, then kept publishing the entries after the failed one and
acked them once the shared .idx sync went through. When the failed put
tore its .idx row, the rows appended after it land off alignment, so the
next load parses them as garbage and the acked writes are gone.

Once a durable entry fails to publish, refuse every later entry of the run
with ReadOnly, as the per-needle path does. The entries before it stay
acked; their rows go down with the run's one .idx sync. The refused
records stay on the .dat unindexed, as the failed one does on its own.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: refuse a grouped entry staged as a cookie mismatch too

After a durable entry in a grouped run fails to index, the entries
after it are refused as they would be on their own. On its own an entry
meets check_writable before its cookie check, so one staged as a cookie
mismatch now gets the refusal too, instead of keeping its staging error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: trim a torn .idx row back so the next stays aligned

A failed write_index_entry can leave half a row in the .idx. With the
writer appending at the tail, every row written after it lands off
alignment and the next load parses them as garbage, so a write acked
behind a torn row does not come back. Trim the file back to
idx_file_offset on a failed append, in both needle maps, and cover it
with a test that writes past a torn row and reloads.

* volume server: refuse queued Go writes once a durable index update fails

processBatch kept writing after a failed nm.Put, and the single-write
path checked IsReadOnly only outside the volume lock. A durable write
whose index update fails now marks the volume noWriteOrDelete, and each
queued request is checked before it appends, so the ones after a failed
durable entry are refused the way a lone write is. Deletes get the same
noWriteOrDelete refusal a lone delete gets.

* volume server: refuse appends while a torn .idx row cannot be trimmed

When trimming back a half-written .idx row itself fails, the next append
would land after the torn bytes and every later row would parse off
alignment on load. Latch the map as torn and refuse appends until the
trim succeeds, on both CompactNeedleMap and RedbNeedleMap; the same
latch covers an orphan row that could not be trimmed after a failed
redb commit.

The .idx writer is now opened with write+append access so truncate_to
(set_len) works on Windows, where an append-only handle cannot trim.

* volume server: write .idx rows at idx_file_offset, not via append mode

Rust's OpenOptions on Windows strips FILE_WRITE_DATA whenever append is
set so the handle stays strictly append-only, which makes set_len fail -
the torn-row trim could never succeed there. Open the .idx writer with
plain write access and seek to idx_file_offset before each row, the same
positioned-write model the Go server uses.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 21:04:30 +08:00
Chris Lu eafe79ebff filer: skip UpdateEntry when inline content is unchanged (#11580)
* filer: skip UpdateEntry when inline content is unchanged

SaveInsideFiler rewrites config files (IAM identities, filer.conf,
remote mappings, policies) unconditionally. Each no-op UpdateEntry is a
metadata event the local meta log persists to /topics/.system/log,
which appends a chunk to a volume. A client that rewrites identical
config on a timer, e.g. the seaweedfs-operator 5-minute resync calling
UpdateUser with unchanged actions, keeps .dat/.idx files growing on an
otherwise idle cluster and prevents HDD spindown
(seaweedfs/seaweedfs#11571).

Skip the UpdateEntry when the stored inline content is byte-identical,
so unchanged writes produce no metadata event and no volume writes.

* filer: test that identical SaveInsideFiler writes skip UpdateEntry

* filer: require stamped Md5 before skipping identical writes

An entry holding identical content but no Md5 (written before hashing,
or by a tool that cleared it) would never get the stamp that
IF_ETAG_MATCH conditional writes key off. Skip only when both the
stored content and its Md5 match, so one write still lands to repair
the stamp.
2026-10-03 21:00:51 +08:00
562afa8ec9 filer: resume metadata subscriber from processed watermark on reconnect (#11574)
* filer: resume metadata subscriber from processed watermark on reconnect

* filer: take the reconnect position from GetResumeTsNs verbatim

The callback is the subscriber's durable resume point; falling back to
StartTsNs when it returns zero can resume from a cursor the log-chunk
reader advanced past still-pending work.

* filer: advance the stream cursor once a retried event recovers

RetryForeverOnError resolves the failure inside handleErr, so returning
without moving StartTsNs replays work the event already did when the
stream reconnects before the next one arrives.

* filer: let filtered-progress markers move the processed watermark

A marker means the source examined everything up to its timestamp and
skipped what did not match the subscription. With a resume callback the
marker now reaches the consumer, and AddSyncJob advances the watermark
to it once every earlier job finished and no failure pins the offset.
Idle filtered stretches no longer rescan on every reconnect, while the
guards keep the watermark behind pending or failed work.

* filer: unpin the watermark once a failed event completes

oldestFailedTsNs was only ever set, so a failure that a replay later
fixed still held the resume offset, and every reconnect re-read the
same backlog. Track outstanding failures in a set and recompute the
pin when the failed event's job finally succeeds.

* filer.remote.gateway: resume bucket sync from the processed watermark

The bucket-sync subscriber runs the same MetadataProcessor queue as
filer.remote.sync; give it the same GetResumeTsNs callback so a
reconnect resumes from durably processed work, not the last seen event.

* util: treat a peer-sent gRPC Canceled as transient

A peer tearing down its end of the transport reports codes.Canceled
("the client connection is closing"), which IsTransientError used to
reject: the sync job then failed on the first try and held the offset
until a restart. Caller's own cancels are still excluded up front by
errors.Is(err, context.Canceled), so only teardown-style statuses take
the new branch.

* fix: preserve filtered progress and distinguish caller cancellation

* filer: bound the failed-event ledger past a persistent outage

A destination rejecting every event grew failedTs by one entry per source
event for the life of the processor. Past maxFailedSyncEvents the set now
collapses to a sticky pin at the smallest failure seen, so the watermark
still replays from the oldest failure while memory stays bounded; a
restart re-derives the exact set.

Also keep a resume-callback consumer's chunk-ref replay filter at the
subscribe-time position instead of option.StartTsNs, so a resubscribe does
not filter out events whose async processing is still pending.

* filer: key the failed-event ledger by event, not just timestamp

A success for one event cleared the pin recorded for a different event
that shared its TsNs, letting the watermark pass an unresolved failure.
The ledger now keys on the event's path identity, so recovery unblocks
only the event that actually failed.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 20:57:50 +08:00
Chris LuandDevin 6c07a5fdd0 s3: keep small ranged GETs on range reads, no whole-chunk downloads (#11577)
* filer: keep a ranged read in random mode through its contiguous tail

A far ReadAt on a fresh ReaderPattern left the sequential counter at -1,
so the next buffer of the same ranged request landed on the frontier and
flipped the verdict straight back to sequential — readChunkSliceAt then
paid a whole-chunk fetch for the remainder of the range. Drop the
counter to -ModeChangeLimit when random mode is entered so the verdict
needs sustained sequential evidence to undo, matching the hysteresis an
established sequential stream already gets.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: pin small ranged GETs to range reads

A ranged GET whose first read lands within SeqTolerance of offset 0 is
judged sequential immediately, and even a far-starting range could flip
back mid-request; either way readChunkSliceAt downloads each covered
chunk in full, multiplying disk reads for small ranged reads (measured
~7x). Pin random mode for ranged requests no larger than SeqTolerance so
all of the request's buffer reads stay range fetches. Larger ranges keep
the dynamic pattern, where whole-chunk fetches amortize.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: fetch only the part of a chunk the view covers

Replaces the PinRandomMode size heuristic with a per-chunk coverage rule.
ViewFromVisibleIntervals already clips chunk views to the request window,
so a view that is not IsFullChunk() is one the request only partially
needs; fetch it as a range regardless of the detected read pattern.

This closes the holes a request-size pin left open: ranges larger than
SeqTolerance no longer revert to whole-chunk downloads once their buffers
look sequential, and ranges that fully cover a chunk keep the shared
whole-chunk path instead of fetching 256KiB slices piecemeal. Prefetch
(MaybeCache) skips clipped views so it cannot amplify a range read either.

PinRandomMode is dropped: no caller needs it once coverage drives the
fetch choice. Range fetches route through fetchChunkDataFn so tests
observe them the same way as whole-chunk downloads.

* filer: keep ciphered chunks on the whole-chunk path

A range fetch cannot save bytes for a ciphered chunk: readEncryptedUrl
always downloads and decrypts the whole blob before slicing. Sending
partial views of ciphered chunks through fetchChunkRange would repeat the
full download per buffer, so they keep the shared whole-chunk path where
one download serves every buffer. Prefetch stays enabled for them for
the same reason.

* filer: keep compressed chunks on the whole-chunk path

Like ciphered chunks, a range request on a compressed chunk makes the
volume server read and decompress the whole needle, so range-per-buffer
would repeat the full backend read for each 256KiB window. Route them
through the shared whole-chunk path via ChunkView.CanRangeFetch.

* filer: fall back to range fetch when a chunk exceeds the reader budget

A ciphered or compressed chunk larger than readerCacheSizeMB can never
be read through the whole-chunk path — the budget rejects the buffer —
so its partial views must still range-fetch or the GET fails outright.

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 15:01:04 +08:00
07da302da0 volume server: ec.decode verifies, cleans up and compacts like Go, off the runtime (#11547)
* volume server: ec.decode reads the .ecx from the index dir it was copied to

VolumeEcShardsCopy writes the .ecx/.ecj into the receiver's -dir.idx, so
with a split data/index dir the decode target has no .ecx beside its
shards. VolumeEcShardsToVolume sized the .dat from the right .ecx but
built the .idx from the data dir, failing with NotFound after the .dat
was already published. It now reads .ecx/.ecj from where the EC volume
opened them and writes the .idx beside the .dat, where Go leaves it.

The live-entry check and the .dat size also ignored deletions recorded
only in the .ecj, which Go folds into the .ecx (RebuildEcxFile) first:
a fully deleted volume was decoded instead of reported as having no live
entries, and deleted tail needles were copied into the .dat. Both now
treat journaled ids as deleted, without rewriting the sealed .ecx.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps the decoded volume writable and reads every .ecj

The rebuilt .idx copied a journaled tail needle's .ecx row verbatim after
the .dat was cut short before it, so the mount saw a row past EOF and
marked the decoded volume read-only. Rows of deleted needles the .dat no
longer holds are now dropped, and each journaled needle still in the .dat
gets one tombstone instead of one per journal entry.

VolumeEcShardsCopy appends journals collected from other holders into
the idx dir, but the decode read only the .ecj beside the .ecx, which
sits in the data dir when this server generated the shards. It now
reads both, once, in bounded chunks via the loader EcVolume uses.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: test ec.decode drops a sealed .ecx tail tombstone

Covers the other half of the rule added in the previous commit: a tail
needle tombstoned in the .ecx itself (Go's RebuildEcxFile) is cut from
the .dat, and its row must not reach the rebuilt .idx either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode runs its file I/O off the async runtime

VolumeEcShardsToVolume released the store lock before decoding, but read
the .ecx/.ecj, rebuilt the .dat and wrote the .idx inside the async
handler, parking a runtime worker for the length of a volume-sized copy.
The decode now runs in spawn_blocking on inputs snapshotted under the
store lock.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode checks the rebuilt .dat is complete

Go stats the decoded .dat before writing the .idx (VerifyDecodedDatFile)
and fails the decode when it is shorter than the extent the EC index
references, since the caller deletes the shards once the call returns.
The Rust handler returned success without that check. The rebuild
already fails on a short shard read, so this guards the published file
itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode drops the decoded volume's bitrot sidecars

Go removes <base>.ecsum and <base>.ecsum.v<N> beside the .dat and beside
the .ecx once the .idx is written, so a stale checksum sidecar cannot
pass for the protection of a later re-encode. The Rust handler left them
in place. Removal is best effort, as in Go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode compacts the decoded volume

Go ends VolumeEcShardsToVolume with an offline CompactVolumeFiles, so the
decoded volume holds only live needles. The Rust decode left every needle
deleted through the .ecj in the .dat, tombstoned in the .idx, until a
later vacuum reclaimed it.

Store::compact_volume_files loads the unmounted volume, checks free space
the way the vacuum does (the estimate now lives in one helper), and runs
the vacuum's compact-by-index and commit. As in Go a failed compaction is
logged and the decode still succeeds, so the uncompacted .idx rules stay:
the tests that pin them now make the compaction fail.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps deletes journaled while the .dat is written

The decode read the .ecj journals once, before rebuilding the .dat, so a
delete that reached the EC volume during the rebuild was left out of the
new .idx and the needle came back live. Each journal's read length is now
kept, and the bytes appended since are read just before the .idx is
written, after waiting out any journal append in flight (appends hold
the store write lock), so every delete acknowledged by then is in the
.idx. A delete after that point is still lost, as in Go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Guard overlapping ec decode requests; serialize journal catch-up

volume_ec_shards_to_volume runs its decode in spawn_blocking, so a
dropped request leaves the job running and a retry would race it on the
temporary and final volume files. Claim the vid in a per-server
in-flight set until the blocking job finishes, and return Unavailable
to an overlapping request. The Go handler has the same exposure and
gets the same guard.

Journal appends hold the store write lock through their
sync-or-truncate, so holding a read lock across the catch-up read
guarantees every record it sees is committed: a rolled-back delete can
no longer leave a tombstone in the decoded index.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Reconcile the swap when offline compaction commit fails

A CommitCompact that fails after the .cpc marker may have renamed .dat
but not .idx. cleanup_compact refuses while the marker exists, so the
mismatched pair survived until a restart reconciled it — and the decode
caller treats the failure as non-fatal. Run reconcileCompactState on
commit failure so a decided swap rolls forward and orphan temps are
removed before the volume can mount.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Release the decode claim on panic

* volume: add ec_decodes_in_flight to the integration-test state literal

* volume server: hold the decode tail's lock through compaction

The catch_up read released before the rebuilt .idx was written and the
volume compacted, so a delete synced to .ecj in that window was durably
journaled yet absent from the published index — resurrecting the needle.
Rust now holds the store read lock from catch_up through compact, and Go
mirrors it by holding the volume's journal lock from the journal-
consuming index write through CompactVolumeFiles.

* volume server: serialize ec decode's tail per volume, not per store

Review follow-ups on the decode path:

- Rust: holding the store read lock from journal catch-up through the
  offline compaction stalled every writer on unrelated volumes for the
  whole rewrite. The new ec_decode_tail set marks the vid only while its
  .idx is published and .cpd/.cpx swapped; the two local .ecj append paths
  (VolumeEcBlobDelete, the distributed delete's local journal) wait on a
  Notify for that span — Go's per-volume ecjFileAccessLock semantics
  without the global stall. VolumeMount and the staged-adopt path are also
  held off while a decode claim is in flight so neither can race the swap.

- Rust: the initial journal read ran unlocked, so bytes a rolled-back
  append later truncated could be folded in as phantom tombstones. The
  first pass stays unlocked (a slow journal must not stall the store) and
  a rescan under the quiescing read lock re-reads only committed content;
  catch_up now rebuilds the id set when a regular journal shrank.

- Go: the decode resolved the compaction DiskLocation through
  FindEcVolume while holding the journal lock, inverting DestroyEcVolume's
  map->journal order into a deadlock. The lookup now happens first, and
  DestroyEcVolume/deleteEcVolumeById/DiskLocation.Close destroy outside
  the map lock.

- Go: RebuildEcxFile unlinks .ecj while the volume's ecjFile handle stays
  open, so later deletes could commit to a detached inode. Both call sites
  now fold under the journal lock and ReopenDeletionJournal repoints the
  handle at the live path, working on the volume's resolved .ecx dir
  (EcIndexBaseFileName) rather than the configured index dir.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: fence EC remounts behind the destroy tombstone

DestroyEcVolume, deleteEcVolumeById, and the collection-delete sweep now
remove the EcVolume from ecVolumes before destroying it off-lock, so a
concurrent remount could re-open shard files that the in-flight destroy
then unlinks — registering a detached fd.

Each destroy records a per-vid tombstone channel in a new
ecVolumesDestroying map before dropping the map entry and closes it when
Destroy returns. The tombstone intentionally survives as the vid's
destroy generation: loadEcShardWithIdxDir compares it before and after
opening the shard, so a destroy that both started and finished inside the
open window is still detected. A mismatch drops the just-opened shard
(releasing its fd and mount gauge) and retries after the destroy
completes; a successful mount clears the stale tombstone.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: rescan the .ecj under the store lock only after a rollback

The decode's second journal pass ran a full rescan under the store read
lock on every decode, stalling unrelated writers for the length of the
scan. Bump a process-wide epoch whenever a failed append truncates its
uncommitted tail; an unchanged epoch between the unlocked read and the
quiesced pass proves every id folded in was committed, so catch_up()
suffices. catch_up() also treats a journal that was read but has since
disappeared as shrunk to zero, so its earlier ids cannot linger.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: check the decode tail under the store write lock on delete

A blob delete waited for the publishing tail before taking the store
write lock, so a decode that claimed the tail while the delete was
parked behind the decoder's read lock could still see the journal append
land after the rebuilt .idx — an acknowledged delete the mount would
miss. Test tail membership under the write lock instead, retrying after
the wait; journal_delete_local reports WouldBlock for the same recheck
on the distributed path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: claim the vid for mount and staged adoption, per volume

VolumeMount and the staged .copying adoption held the
ec_decodes_in_flight set lock through slow file renames and mounts,
stalling every unrelated volume's decode, mount, and adoption. Take the
per-volume claim instead — the same exclusion against a racing decode
for this vid, released when the call returns.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: fail the decode when a compaction commit marker survives

CompactVolumeFiles' caller logged a compaction error and went on to
delete the EC shards. When the commit marker (.cpc) is still on disk the
.dat/.idx swap was decided but could not be reconciled, so the mounted
pair may be mismatched — report the failure instead so the shards are
kept and the caller can retry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: gate the parked-delete test on the held write lock

The releaser thread and the spawned delete raced for the store write
lock; on a slow runner the delete could acquire it first and commit
before the tail was ever claimed, failing !delete.is_finished() on the
Windows unit-test job. Spawn the delete only after the thread reports
the lock held.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:55:15 +08:00
68df7511f6 filer.remote.sync: do not pin the sync offset on completed work (#11569)
* filer.remote.sync: do not pin the sync offset on completed work

* filer.remote.sync: a superseded rename uploads the current entry; typed NotFound for a stamp on a deleted entry

* filer.remote.sync: a superseded rename keeps the old key when it is the only copy and uploads once

* filer.remote.sync: a rename whose content is now remote-only fails the event instead of completing it

* filer.remote.sync: a remote-only rename copies the old object to the destination before deleting it

* filer.remote.sync: the remote-only rename path follows the filer's current entry and verifies the destination object

* filer.remote.sync: an event that described an entry without data is superseded once the filer wrote to it

* filer.remote.sync: a superseded rename does only the work left to do

uploadCurrentEntry met a remote-only current entry with a fixed error, but a
sync plus remote.uncache in the meantime leaves the destination holding the
stamped object; that state is complete, not lost. The remote-only case now
finishes through completeRemoteOnlyRename, which verifies the destination
against the entry stamp and fails only when neither key holds the content.

A current entry whose stamp covers its content was already uploaded by the
superseding event; skip it instead of writing the same bytes again.

* filer.remote.sync: an inherited stamp does not prove the content synced

The stamp-coverage skip in uploadCurrentEntry read LastLocalSyncTsNs as
proof the current content was uploaded, but a rename carries the source
entry's stamp to the destination: a rewrite hidden by that stamp (the case
the fallback upload exists for) carries a LastLocalSyncTsNs at or after its
mtime and would have been skipped. Drop the check; the remote-only path
verifies content at the destination itself through describes.

---------

Co-authored-by: James Sas <james@medable.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 12:34:52 +08:00
Chris Lu 793ce06b10 s3api: allow unsigned SSE-C customer key headers on presigned requests (#11578)
* s3api: allow unsigned SSE-C customer key headers on presigned requests

AWS requires only x-amz-server-side-encryption-customer-algorithm to be signed on presigned URLs; the key and key-MD5 headers are supplied at request time. Since #9121 rejected any x-amz-* header outside SignedHeaders, SDK-generated presigned SSE-C requests (e.g. .NET GetPreSignedUrlRequest) fail with SignatureDoesNotMatch. Exempt the customer key and copy-source key headers for presigned requests only.

* s3api: test presigned SSE-C requests carrying unsigned key headers
2026-10-03 12:33:52 +08:00