Commit Graph
10000 Commits
Author SHA1 Message Date
68df7511f6 filer.remote.sync: do not pin the sync offset on completed work (#11569)
* filer.remote.sync: do not pin the sync offset on completed work

* filer.remote.sync: a superseded rename uploads the current entry; typed NotFound for a stamp on a deleted entry

* filer.remote.sync: a superseded rename keeps the old key when it is the only copy and uploads once

* filer.remote.sync: a rename whose content is now remote-only fails the event instead of completing it

* filer.remote.sync: a remote-only rename copies the old object to the destination before deleting it

* filer.remote.sync: the remote-only rename path follows the filer's current entry and verifies the destination object

* filer.remote.sync: an event that described an entry without data is superseded once the filer wrote to it

* filer.remote.sync: a superseded rename does only the work left to do

uploadCurrentEntry met a remote-only current entry with a fixed error, but a
sync plus remote.uncache in the meantime leaves the destination holding the
stamped object; that state is complete, not lost. The remote-only case now
finishes through completeRemoteOnlyRename, which verifies the destination
against the entry stamp and fails only when neither key holds the content.

A current entry whose stamp covers its content was already uploaded by the
superseding event; skip it instead of writing the same bytes again.

* filer.remote.sync: an inherited stamp does not prove the content synced

The stamp-coverage skip in uploadCurrentEntry read LastLocalSyncTsNs as
proof the current content was uploaded, but a rename carries the source
entry's stamp to the destination: a rewrite hidden by that stamp (the case
the fallback upload exists for) carries a LastLocalSyncTsNs at or after its
mtime and would have been skipped. Drop the check; the remote-only path
verifies content at the destination itself through describes.

---------

Co-authored-by: James Sas <james@medable.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 12:34:52 +08:00
Chris Lu 793ce06b10 s3api: allow unsigned SSE-C customer key headers on presigned requests (#11578)
* s3api: allow unsigned SSE-C customer key headers on presigned requests

AWS requires only x-amz-server-side-encryption-customer-algorithm to be signed on presigned URLs; the key and key-MD5 headers are supplied at request time. Since #9121 rejected any x-amz-* header outside SignedHeaders, SDK-generated presigned SSE-C requests (e.g. .NET GetPreSignedUrlRequest) fail with SignatureDoesNotMatch. Exempt the customer key and copy-source key headers for presigned requests only.

* s3api: test presigned SSE-C requests carrying unsigned key headers
2026-10-03 12:33:52 +08:00
Chris Lu 0ca484c354 vacuum: bound master vacuum RPCs with phase deadlines (#11579)
* vacuum: bound the commit RPC with a phase deadline

VacuumVolumeCommit ran on context.Background(), so a volume server that
keeps the call pending would hold the topology-wide vacuum guard
forever and every later sweep would be skipped. Give the call a
deadline scaled like the existing phase waits (one minute per GB of
the volume size limit) so a stalled commit ends as an error instead of
blocking the sweep; the timeout is a var so tests can shrink it.

* vacuum: bound the replica status probe with a phase deadline

The VolumeStatus call on replicas that were not compacted also ran on
context.Background(), so a stalled replica could pin the sweep the
same way a stalled commit can. Give it the same per-phase deadline.

* vacuum: bound the cleanup RPC with a phase deadline

VacuumVolumeCleanup also ran on context.Background(); a stalled
server would keep the sweep worker and the shared vacuum guard
pending forever. Give it the same per-phase deadline.

* vacuum: let the check and compact phase waits cancel their RPCs

The coordinator wait timers fired while the check and compact calls
still ran on context.Background(), so the sweep gave up but the RPC
goroutine stayed until the server answered, and a compact stream kept
writing on the server. Share one deadline context between the wait and
the calls so an expired wait actually cancels them.

* vacuum: test that a stalled volume server releases the vacuum guard

A fake volume server keeps one vacuum-phase RPC pending until the
client context is cancelled. Before the phase deadlines, Vacuum never
returned and vacuumLockCounter stayed held; now each phase cancels on
its deadline and the guard is free for the next request.

* volume: stop compaction at the next needle when the client cancels

The progress callback only noticed a gone client when a 128 MiB report
failed to send, so an aborted VacuumVolumeCompact kept copying for up
to a whole interval while the master had already moved on to cleanup.
Check the stream context on every needle, the same early return the
Rust volume server does with tx.is_closed().

* vacuum: assert the stalled phase RPC is cancelled, not just bypassed

The check and compact coordinator waits already returned on timeout
before the deadlines existed, so a regression that put the calls back
on context.Background() would pass unnoticed. Wait for the fake server
to report that the phase RPC context ended.

* vacuum: give the stalled-RPC test room to reach the handler

The 50ms phase budget starts before goroutine scheduling and the gRPC
dial, so a busy test host could expire it before the fake server saw
the call. Raise the override to 250ms; the test still finishes in
about a second.

* vacuum: describe the phase deadline as scaled, not per-GB

The formula keeps the exact expression the check and compact waits
already used (floor plus one at 1 GiB granularity); it is a backstop,
not a per-GB SLO.
2026-10-03 12:32:42 +08:00
yi111andYi-111-a 3e679e925e filer: keep generated inodes inside the positive signed 64-bit range (#11567)
AsInode derives inodes from HashStringToLong, which is uniform over int64,
so roughly half of the derived values land above math.MaxInt64 once they are
converted to uint64. The Elasticsearch store indexes Entry.Attr.Inode as a
signed long, so those values are rejected with HTTP 400 and the metadata
entry is never written, which the filer then retries forever.

Fold the sign bit off in one place, util.NormalizeInode, and route both
derivation sites through it: FullPath.AsInode (path plus creation time) and
the hard-link branch in ensureEntryInode (HardLinkId hash). Masking keeps
the other 63 hash bits, so distinct paths still get distinct inodes, and it
applies identically to the FUSE mount, which derives the same value.

Co-authored-by: Yi-111-a <34116709+0-xiaosu@users.noreply.github.com>
2026-10-03 09:13:59 +08:00
Peter DoddandChris Lu 52fb9f93ff s3: track filer joins and leaves pushed by the master (#11563)
* fix(s3): track filer joins and leaves pushed by the master

The S3 FilerClient replaced its -filer seed with a master snapshot of
filer IPs at boot and refreshed it only every 5 minutes. A rolling
restart replaces every filer well inside that window, leaving S3
servers with only dead addresses and failing every write until the
next poll.

Apply the master's ClusterNodeUpdate pushes to the filer list as they
arrive, keeping the poll as a backstop. The last filer is never
removed, and a poll snapshot requested before a push was applied is
discarded rather than overwriting newer membership.

* Defer last-filer leaves; bump the generation only on real changes

* fix(s3): cancel deferred filer leaves on rejoin and on discovery

A deferred last-filer leave outlived the filer it was recorded for: a
rejoin at the same address looked like a duplicate add, and a discovery
snapshot left the entry behind. The next join then removed a live
filer until the following poll.

A join now cancels any deferred leave for its address, and an applied
snapshot clears them, since it is the master's current membership.

* Bump the push generation when a rejoin cancels a deferred leave

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-03 09:10:43 +08:00
0ff7794c54 volume: compact an oversized .ecj at mount, safely (Rust + Go) (#11555)
* volume: compact an oversized .ecj at mount, safely (Rust + Go)

Restore the mount-time compaction dropped from #11408, Rust + Go parity.
A journal already bloated by repeated shard copies is folded down to the
id set it encodes.

- Trigger after load when file_records > max(threshold, 4x distinct),
  with a 1 MiB floor so small journals are never rewritten. The set is
  written to .ecj.compact.tmp + fsync, the handle dropped, renamed,
  the directory fsynced and the append handle reopened. A failure before
  the rename keeps the original journal and handle; a failure after it
  fails the mount.
- Go never compacts after a failed journal load; the set would be
  partial and the rewrite would drop the unread records.
- A per-path registry (ecj_registry.rs / ecj_registry.go) counts EcVolume
  holders and out-of-band writers of each .ecj. Compaction runs only
  when this volume is the sole holder and no copy is writing; holders
  and writers wait while one runs. This covers shared -dir.idx journals
  and cross-disk reconcile, where another EcVolume may hold the same
  journal.
- VolumeEcShardsCopy and EC index recovery register as writers around
  their .ecj append and partial-file cleanup.
- Under the reservation, re-check that the file on disk is still the
  inode and size that was loaded.
- Publish errors are classified where they happen; a failed rename plus
  a failed restore reports both errors.
- Compaction runs after the .vif / bitrot checks, so a refused mount
  leaves the journal untouched.
- The tmp is opened like other volume files, removed at mount if a crash
  left it, and listed in every EC index cleanup path.

Failure paths are tested through the real mount via injectable fs steps
(open_with / newEcVolumeWith), plus sibling holders, active copies,
changed-after-load, stale tmp cleanup, refused mounts and the Go
load-error guard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: fail the mount when the compacted .ecj's directory cannot be synced

The Rust mount synced the journal's directory after renaming the compacted
file over it through the crate's best-effort fsync_dir, which returns Ok
when the directory cannot be opened. A rename needs only write and search
permission, so on a directory without read permission the replacement was
published, never synced, and the mount went on taking deletes against it.

Sync through a helper that propagates the open error, as Go's
util.FsyncDir already does, so that case fails the mount like any other
post-rename sync failure.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: test the no-compaction-after-failed-load rule through the Go mount

The test for it handed compactEcjAfterLoad an artificial error on a volume
that had loaded cleanly, so it would not notice NewEcVolume dropping the
real load error on the way to compaction.

Make the journal read one of the injectable ecjFsOps steps and fail it
inside the real mount, after the first chunk, on a journal whose last
entry is an id the first chunk does not hold. The mount must leave the
file byte for byte as it was; a clean remount then compacts and keeps
that id. The Rust mount fails outright on a load error, so it has no
equivalent path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: register ReceiveFile's .ecj writes with the journal registry

ReceiveFile refuses a mounted EC volume only once, when the info message
arrives, then creates the .ecj and streams chunks into it. A volume that
mounted on that journal mid-stream could find a bloated prefix, pass the
inode-and-size re-check and rename a compacted file over it; the rest of
the stream then went to the unlinked inode and was lost.

Register the path as a writer before the file is created, in both the Go
and Rust handlers, and hold it until the file is closed and any partial
copy removed, as the shard-copy and index-recovery appends already do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: skip .ecj compaction when a writer ran since the journal was loaded

Compaction checked only that no writer was active at the reservation, and
that the file was still the loaded inode at the loaded size. A ReceiveFile
truncates and refills the journal in place, so one that ran during the
mount's load, or after it, and finished before the reservation could leave
different ids at the same length; compaction then wrote the stale set over
them.

Give each path a write generation that every writer bumps as it starts. A
holder records it, and whether a writer was active, when it registers,
which is before it opens and loads the journal. It may compact only if no
writer was active then and the generation has not moved. Same rule in Go
and Rust; the journal read becomes an injectable step in Rust as it is in
Go, so both test the in-place rewrite through the real mount.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: match the ReadOnly(VolumeId) variant in write_volume_needles

#11543 matched VolumeError::ReadOnly as a unit variant in Store::write_volume_needles, and #11544 changed it to ReadOnly(VolumeId) in the same merge window. Each passed CI on its own, but master no longer compiles the Rust volume server. Carry the volume id through.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 09:08:19 +08:00
30069f3e45 iam: manage OIDC providers and roles over the filer IAM gRPC service (#11523)
* s3/iam: manage roles through the IAM API, with an opt-in persistent role store

Roles could only come from the IAM config file: the S3 server pinned the
role store to memory and the embedded IAM API had no role actions, so a
role could not be created, retrusted or revoked without editing the file
and restarting every gateway.

Role store
- Read the `roleStore` key (the IAMConfig field already existed). With an
  IAM config file the default stays memory; with none it is the filer, as
  for OIDC providers, so zero-config clusters keep runtime-created roles.
- Roles from the IAM config file never go into a persistent role store,
  which outlives the file and may be shared by S3 servers with different
  files. They are served from memory beneath the store, as OIDC providers
  are: a stored role of the same name takes precedence, and deleting it
  restores the file's. A config-file role cannot be changed or deleted
  through the API (UnmodifiableEntity), and removing one from the file
  removes it at the next start. An in-memory store holds them as records,
  as before. They have no creation time, so CreateDate is omitted rather
  than reporting when this server started. SetRoleStore installs a store
  the same way, so a store set after startup keeps the config-file roles,
  as SetOIDCProviderStore does for providers.
- Watch /etc/iam/roles and drop the cached role definitions on change. The
  cached filer store otherwise serves a peer's stale role for up to its 5m
  TTL, which keeps a revoked trust policy in force on the other gateways.
- Role stores wrap ErrRoleNotFound for a missing role; the filer store
  used to report any failed lookup as "role not found". CreateRole proceeds
  only on a confirmed absence, so an unreadable store cannot let it write
  over an existing role.

IAM actions
- CreateRole, GetRole, ListRoles, DeleteRole, UpdateAssumeRolePolicy,
  AttachRolePolicy, DetachRolePolicy, ListAttachedRolePolicies. The reads
  are allowed in read-only mode.
- A role defined in the config file is reloaded from it at every start, so
  changing or deleting it through the API is refused (UnmodifiableEntity)
  rather than silently reverted.
- DeleteRole with policies attached is refused (DeleteConflict), as on AWS.
- Role names follow AWS's rules ([\w+=,.@-]{1,64}); a role is stored as
  <name>.json in the filer, so this also keeps a name from leaving the role
  store's directory. At most 10 managed policies per role (AWS's default
  quota; MaxManagedPoliciesPerUser is 10 too), LimitExceeded beyond.
- DeletePolicy is refused (DeleteConflict) while a role attaches the
  policy, as it already is for users and groups: roles attach policies by
  name, so a policy created later under the deleted one's name would
  otherwise take effect on the role.
- Role paths other than "/" and role tags are not stored, so they are
  refused rather than dropped.

Role IDs and sessions
- Roles get a unique RoleId when first stored (random, AWS AROA form),
  kept across updates; a config-file role gets a stable ID derived from its
  name, since it is created again at every start.
- Sessions issued through AssumeRoleWithWebIdentity, AssumeRoleWithCredentials
  and AssumeRole carry the role's ID (claim "rid"), and a request under a role
  whose current ID differs is denied. Resolving a session's policies by role
  name let a session outlive its role: once a role was deleted, a role later
  created under the same name — with a different trust policy and different
  policies — revived every unexpired session of the old one with the new
  role's permissions. Sessions issued before this change carry no ID and are
  unaffected until they expire.

Integration test (test/s3/iam, run with `make start-services`):
TestWebIdentityWithProviderAndRoleManagedThroughIAMAPI configures an OIDC
provider, a managed policy and a role entirely through the IAM API against a
JWKS served by the test, then checks the trusted subject gets credentials
scoped to the attached policy; another subject, a token signed by another
key, an unsigned token and a token for another audience are refused; and UpdateAssumeRolePolicy moves the
trust at once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: manage OIDC providers and roles over the filer IAM gRPC service

The filer's SeaweedIdentityAccessManagement service covers users, access
keys, policies and service accounts, but not the OIDC providers and roles
that STS web-identity federation needs. A controller that already manages
IAM over this service (seaweedfs-operator's S3OIDCProvider) has no
transport for them; its swadmin client returns ErrOIDCNotWired and names
this as the recommended fix.

- PutOIDCProvider / GetOIDCProvider / DeleteOIDCProvider / ListOIDCProviders
  and PutRole / GetRole / DeleteRole / ListRoles.
- They write the filer-backed stores at their default paths, which S3
  servers read when configured with a filer-typed "oidcProviderStore" and
  "roleStore"; the S3 servers' /etc/iam subscription applies changes
  without a restart.
- Put is an upsert, so a controller can reconcile to it. Deleting a
  provider or role that does not exist returns NotFound, as DeleteUser does
  for a user; clients treat that as already deleted. The provider's account
  ID travels in the request, since the filer does not know the STS
  accountId.
- PutRole applies the IAM API's rules: AWS role names, at most 10 managed
  policies.
- An S3 server serves the roles and providers of its own IAM config file
  ahead of the store, so a stored entry with the same name has no effect
  on that server.
- PutRole keeps a replaced role's RoleId and gives a role created anew a
  fresh one, so sessions of a deleted role do not carry over to a later role
  of the same name.
- DeletePolicy returns FailedPrecondition while a role attaches the policy
  (see the IAM API's DeleteConflict in the previous change). DeletePolicy on
  this service still does not check user attachments, which predates this.
- PutOIDCProvider requires an https issuer (http only for a loopback host):
  STS fetches the issuer's signing keys from it, so over plain HTTP anyone
  on the network path could substitute their own.
- The OIDC provider and role RPCs refuse to run on an unauthenticated
  service (FailedPrecondition until jwt.filer_signing.key is set). Users and
  policies keep the service's opt-in auth, but these grant STS access
  outright: otherwise anyone who can reach the port could register an issuer
  they control, create a role trusting it, and exchange a token for S3
  credentials. The filer's unauthenticated notice becomes a warning that says
  so.
- A store that cannot be read is Unavailable, never "not found", so a Put
  never writes over an entry it could not see.
- Validation is shared with the IAM API through PrepareRoleDefinition and
  PrepareOIDCProviderRecord.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: bind every role session to its role, and change roles atomically

Review follow-ups.

Session binding
- The role-ID check ran only when a session carried no policy names, and
  AssumeRole embeds the role's attached policies, so those sessions kept
  their permissions after the role was deleted or recreated. The check
  now runs for every session carrying a role ID, before policy selection.
- A named role that cannot be resolved at issuance gets no session,
  instead of one with no role ID (which nothing binds).
- A config-file role's ID is derived from its name and trust policy, not
  the name alone: a different role put in the file under the same name
  gets a new ID, while an unchanged role keeps its sessions across restarts.

Role writes
- RoleStore gains UpdateRole, a read-modify-write that lands only if the
  role is unchanged since the read, and otherwise re-reads and retries. The
  filer store uses the filer's write conditions (IF_NOT_EXISTS for a new
  role, IF_ENTRY_EQUAL otherwise). CreateRole, UpdateAssumeRolePolicy and
  Attach/DetachRolePolicy all go through it, so two gateways no longer
  overwrite each other's changes, a change racing a delete no longer
  writes the role back, and of two concurrent creates one gets
  EntityAlreadyExists.
- The filer store's ListRoles pages past 1,000 entries and fails on a
  broken stream instead of returning what arrived, so DeletePolicy's
  attachment check sees every role. ListRoles skips a role deleted between
  listing and reading it.
- CreateRole validates first; a failed write is ServiceFailure, not
  InvalidInput. Any Tags.* parameter is refused, not only the first key.
- ExecuteAction's skipPersist covers the S3ApiConfiguration only; the
  comment now says so. Role and OIDC provider actions write their own stores.

Each fix has a test that fails without it. Against a real filer with two
gateways, concurrent AttachRolePolicy calls lost 1-4 of 8 attachments per
run before this change and none after.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: PutRole changes roles atomically and checks its ARN; https issuers' keys stay on https

Review follow-ups on top of the role-store changes.

- PutRole goes through RoleStore.UpdateRole, so the decision to keep an
  existing role's ID or mint a new one is made against the role as it is
  when written. A PutRole racing a DeleteRole can no longer write the
  deleted role back with its old ID, which would revive its sessions. A
  failed store read or write is Unavailable.
- PutRole refuses a role_arn that does not name the role: STS resolves a
  role by the name in the ARN it is given.
- PutOIDCProvider requires an https issuer, but discovery could still name
  a plain-http jwks_uri, and a key fetch could be redirected to http. For
  an https issuer, a non-https jwks_uri from discovery is refused (the
  issuer's own /.well-known/jwks.json is used instead), and the client
  that fetches discovery and keys refuses any https-to-http redirect. An
  operator-set jwksUri is left as configured.

Each has a test that fails without its guard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: one role snapshot per decision; DeleteRole is atomic; watch a custom role store path

Review follow-ups.

- Authorization evaluates the policies of the role definition the session's
  binding was checked against, instead of reading the role again: a role
  replaced in between cannot lend a session its policies.
- AssumeRole and AssumeRoleWithLDAPIdentity issue the session from the
  definition whose trust admits the caller (IAMManager.ResolveRoleForPrincipal),
  and take its ID, duration cap and embedded policies from that same
  definition. A role replaced after the caller's trust check by one that does
  not trust the caller now yields AccessDenied, not a session bound to the
  replacement.
- A RoleUpdate that returns nil deletes the role, on the same condition as a
  write: the filer store deletes with ObjectTransaction on IF_ENTRY_EQUAL,
  routed and locked like the conditional CreateEntry. DeleteRole decides
  against the role it deletes, so a policy attached meanwhile on another
  server is a DeleteConflict, and a delete never removes a role written
  after its check.
- S3 servers watch the role store's configured basePath, not only
  /etc/iam/roles, so a custom path also drops peers' cached roles on change.

Each has a test that fails without it. Live against a real filer: DeleteRole
refuses while a policy is attached and removes the entry once detached; all
test/s3/iam CI stages pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: state which roles DeletePolicy's attachment check can see

RolesAttachingPolicy sees the stored roles and this server's config-file
roles. A role defined only in another server's IAM config file is invisible
to it, so a config-file role that attaches a managed policy is protected
only on the servers whose file defines it. The doc comment now says so and
how to avoid it: keep such roles in every server's file, or attach only
config-file policies to config-file roles.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: note that a role store set after startup is not watched for peer changes

S3 servers build their metadata watch list once, at startup, from the role
store installed then. SetRoleStore's doc now says that a filer-backed store
installed later with a different basePath is not watched, so peers' changes
to it reach this server's cached roles only when the cache expires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: DeleteRole deletes only the role it saw; issuer URLs are bare

Review follow-ups.

- The filer IAM service's DeleteRole looked the role up, then deleted by
  name, so a PutRole landing in between had its new definition deleted. It
  now deletes through RoleStore.UpdateRole, conditional on the entry it
  read. If the role was replaced meanwhile, it returns Aborted rather than
  deleting the replacement, and the caller decides again.
- PutOIDCProvider refuses an issuer URL with userinfo, a query or a
  fragment. The provider's ARN comes from host and path alone, while STS
  matches a token's iss claim against the stored URL exactly, so such a
  provider shared the bare issuer's ARN and matched no token. A loopback
  "localhost" is now matched without regard to case.

Both have tests that fail without them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: write OIDC providers atomically over the filer IAM gRPC service

PutOIDCProvider read the record, then stored unconditionally; a racing
DeleteOIDCProvider left the put's stale read merged into the rewritten
record. DeleteOIDCProvider read, then deleted unconditionally; a racing
PutOIDCProvider's newer record could be removed instead. These are the
races the role RPCs closed with UpdateRole.

OIDCProviderStore gains UpdateProvider with the same contract: memory
under its lock, filer as a conditional write (IF_ENTRY_EQUAL /
IF_NOT_EXISTS) or conditional delete retrying a changed entry.
PutOIDCProvider merges the fields the request cannot carry against the
record as it is written; DeleteOIDCProvider aborts rather than delete a
record replaced meanwhile.

isRoleWriteConflict is renamed isEntryWriteConflict — the conditional-
write check is shared by both stores now.

* iam: guard PutRole against a nil credential manager, fix its doc comment

PutRole read attached policies through s.credentialManager without the
nil check its sibling handlers make, so a server built without one
panicked on a PutRole naming a policy. It now fails the call as
FailedPrecondition like the others.

The doc comment also had the store/static precedence backwards: a stored
role shadows a same-named config-file role (as the overlay serves it),
not the other way around.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-03 08:36:19 +08:00
fa77cde7da vacuum: check compaction space against live bytes, not volume size (#11524)
* vacuum: size the compaction space check by live bytes, not volume size

ensureCompactVolumeSpace required the volume's current .dat and .idx size as
free space before compacting. That is the size of the garbage, not of what
compaction writes, so on a disk that filled up until its volumes went
read-only every compaction was refused, including all-garbage volumes that
would compact to a superblock and an empty index. The sweep then retried
every volume each cycle and reclaimed nothing (issue #11516).

Estimate the output from what the needle map already tracks: live content
bytes plus a per-needle framing upper bound behind a superblock, and one
index entry per live needle. The estimate never exceeds the current volume
size and preallocate still wins when larger. Volumes whose deleted sizes are
unknown (.sdx converted back to .idx) keep the whole volume as the estimate.

The disk probe moves behind a package variable so the tests can stand in
for a full disk; the tests build real volumes instead of re-implementing
the formula.

* vacuum: space check reserves the index on top of preallocate, checks a separate index disk

Review follow-ups: preallocate only stands in for the new .dat, so the
rebuilt index is added on top of it; with separate index directories the
data disk is checked for the .cpd and the index disk for the .cpx; and the
estimates carry 1/16 headroom because counters rebuilt from an index file
pass through a Bloom filter with a 0.1% false positive rate. Neither
estimate exceeds the current file.

* vacuum: split the space check by filesystem, not by directory name

Two directories can sit on one filesystem and share its free space, so
the data and index estimates are checked separately only when the index
directory is on another device; otherwise the sum must fit. Unknown is
treated as shared.

* vacuum: ask the index directory for its share even when it looks like the same filesystem

A volume mounted under the data directory's drive letter on Windows has
the same volume name, so the identity check calls it shared. Checking the
index directory for the index estimate as well costs one statfs and
catches a full index mount either way.

* vacuum: identify a Windows volume by its GUID, not its path prefix

A volume can be reached through a drive letter and through a folder it is
mounted on, so filepath.VolumeName says nothing about the free-space pool.
Resolve each directory to its mount point and compare the volume GUIDs;
when that fails the two are treated as shared.

* vacuum: keep the framing and disk_space_low coverage the rebase displaced

* rust volume: split the compaction space check across data and index disks

Mirror the Go check: estimate the new .dat and rebuilt .idx separately —
live content plus per-needle framing capped at the current file, with
preallocate standing in for the data file when larger — and check each
directory against its own filesystem's free space. Two directories on one
filesystem are asked for the sum.

* vacuum: tighten comments on the compaction space check

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 08:31:56 +08:00
35b090a4df volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go) (#11554)
* volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go)

An EC volume's deletion journal is a set of needle ids, but shard copy
and index recovery appended the peer's whole journal, doubling the file
on every ec_balance round trip. Fold the peer's ids in as a union
instead: only ids the local journal lacks are appended.

- The journal is never replaced. A mounted EcVolume merges a peer's ids
  through its live handle under the lock deletes take (Go
  MergeJournal / Rust merge_journal), wherever its journal lives.
- An unmounted journal gets only the missing ids appended while mounts
  are excluded; the delta is read outside the lock and re-read if the
  journal changed.
- The source .ecj streams into memory as an id set: no staging files,
  chunked reads, memory proportional to distinct ids.
- Go and Rust agree that a source journal exists when it sends a
  modified time or any bytes. A missing source stays a no-op.
- Rust runs every merge in spawn_blocking and shares one receive/merge
  path between shard copy and index recovery.

The decode path and the journal format are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: route .ecj merges to the runtime that holds the journal open

Disks sharing one index directory all resolved as the journal's owner, so
the last one won and a sibling's mounted runtime was skipped: the merge
appended behind its open handle and the sibling kept serving the peer's
deleted needles until remount. Callers now name the receiving disk by its
data directory; the merge goes through that disk's runtime, else a
sibling runtime whose journal is the target file.

In Go the unmounted append now holds every disk's EC lock (in location
order) while it rechecks for a mount, so a sibling mounting from this
disk's index during the unlocked read is merged through instead.

In Rust a mount that lands during the read is merged through directly and
its added count returned, rather than discarded and reported as zero.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: sync merged .ecj records outside the disks' EC locks

The unmounted merge held every disk's EC read lock across its fsync, so a
slow sync on one disk held off mounts on all of them, along with the EC
reads queued behind those mounts. Mounts only need to be excluded while
the records are written: the write now happens under the locks and the
fsync after they are released, since a later mount reads the written
records from the page cache. A failed fsync rolls back only if nothing
has mounted the journal or appended to it since the write.

A merge through a mounted volume now keeps only that volume's disk locked
across its fsync.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: roll back an unsynced .ecj merge through a volume mounted mid-sync

If a volume mounted after the unmounted merge wrote its records but before
the fsync failed, the rollback kept the records because the journal was now
open, leaving ids in the volume's deleted set that may never reach disk; a
retried merge then saw them and synced nothing. The rollback now goes
through that volume the way its own failed journal fsync does: truncate
back and drop the ids from the in-memory set, so a retry appends and syncs
them again. It still keeps the records if the volume journaled since, as
truncating would lose that delete. No fsync runs under the disk locks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: decide .ecj merge rollback from the journal's actual length

Two runtimes can hold one journal (cross-disk mounts). The rollback of an
unsynced merge checked one runtime's cached ecjFileSize, which another
runtime's appends leave stale, so it could truncate a delete that runtime
had already synced. The rollback now holds every holder's journal lock and
truncates only if the file's actual length is still the append's end,
then updates each holder's size and deleted set. Otherwise later records
follow the merged ones, so they stay and are rewritten in place and
synced outside the locks, rather than left possibly not durable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: keep unsynced .ecj merge ids out of mounted deleted sets

When a merge's fsync failed, later records blocked the rollback, and the
rewrite-and-sync failed as well, the merged ids stayed in every mounted
volume's deleted set without being shown durable, so a retried merge saw
them as present and synced nothing. They now leave those sets while the
records stay in the file, matching DeleteNeedleFromEcx, which publishes an
id only after its record syncs. The merge returns the error and a retry
appends and syncs them again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: publish merged .ecj ids to every holder of the journal

Two runtimes can journal into the same file when disks share an index
directory. The merge went through only the first holder, leaving a
sibling's in-memory deleted set without the ids, so it could keep
serving a needle the peer deleted until it remounted. Every holder of
the journal now gets the merged ids, in Go and in the volume server.

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Publish merged .ecj ids to the journal actually written

mountedEcJournal prefers the receiving disk's own runtime for the vid,
whose journal may live in its data directory while the copied records
name a sibling's journal in the index directory. Publishing by the
requested ecjPath then marked a holder of a different file deleted on
records that file never persisted, resurrecting the needles on remount.
Publish by the picked runtime's journal path instead.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:46:36 +08:00
1445960f8c filer: persist pending chunk deletions across restarts (durable deletion ledger) (#11550)
* fix(filer): persist pending chunk deletions across restarts

The in-memory FileIdDeletionQueue and DeletionRetryQueue lose every
queued-but-unconfirmed deletion when the filer process restarts. Because
deletions only enter the pipeline through that queue, a crash between
enqueue and the volume confirming the delete leaks the chunk permanently:
nothing remembers it. In a multi-filer deployment this was observed as
growing collections of orphaned chunks after filer restarts, and — via
meta-replay from a peer that still had the entry — orphans being
"resurrected" as live references on the recovered filer.

This implements the "periodic snapshot with recovery on startup" option
noted in the existing DeletionRetryQueue TODO, using the store's KV layer
(no new iterator API required across the 15+ store backends):

- queueDeletions() is the single entry point that keeps the hot in-memory
  queue and the durable ledger in sync.
- Only terminal outcomes (success / not-found / permanent) remove an id
  from the ledger; retryable failures keep it, which is the point.
- A timer and Shutdown() snapshot the pending set to a single KV key.
- On startup, reloadDeletionLedger() re-queues recovered ids after a
  grace window so the initial peer meta-aggregation settles first. This
  avoids a new hazard: purging a chunk that a lagging peer is about to
  re-reference as live data (stale replay turns a stale read into a
  dangling read otherwise).
- Volume deletes are idempotent (not-found == success), so re-deleting
  after a crash never double-frees.
- Kill switch via viper: filer.deleteQueue.persist=false opts out entirely
  (reload also refuses to recover so a stale ledger never comes back).
  Tunables: filer.deleteQueue.persistInterval, .recoveryGrace.

Adds unit tests covering snapshot+recover, retry-keeps-entry, disabled
switch, and zero-value Filer safety (run green under -race).

Co-Authored-By: Athena 🏛️ <hermes-agent@local> (custom / Qwen3.8-Flash-Next-ROCmFP4)

* filer: harden the deletion ledger

- Scope the ledger key by filer address so filers sharing one store do
  not overwrite each other's pending sets; ledgers written under the
  old unscoped key are claimed once on startup.
- Serialize snapshots on deletionSnapshotLock so an in-flight timer
  snapshot cannot overwrite a newer shutdown snapshot, and wake the
  snapshotter on every queue/forget so a queued id persists within
  milliseconds instead of a full interval.
- Merge recovered ids into the pending set immediately on reload; only
  the queue push waits out the grace window, so an early snapshot
  rewrites the recovered ids rather than dropping them.
- A failed or unparseable ledger read blocks persistence for the run
  instead of letting snapshots overwrite the unread ledger.
- Split the ledger into part keys when it exceeds one 64KB value so
  stores with a size cap (FoundationDB) do not strand the backlog.
- GetReadyItems reports retry-exhausted ids so they are forgotten in
  the ledger instead of replaying after every restart.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: close the remaining deletion-ledger durability gaps

- A manifest referencing a missing part is corruption: surface a wrapped
  error and block persistence instead of treating the ledger as absent.
- Multipart snapshots write generation-scoped part keys and publish the
  manifest last, so a crash never mixes old and new part contents.
- Orphaned parts are tracked in a persisted .stale sidecar and retried.
- Legacy/index ledgers are republished under the scoped key before the
  old keys are removed.
- A ledger index lets a filer restart under a new address claim the
  ledger its previous incarnation left behind.
- Expired and permanently-failed retry items only forget the ledger
  epoch they recorded, so they cannot erase a re-queued id.
- A failed startup read no longer disables persistence: every snapshot
  retries the reload until the store reads again.

* filer: tighten ledger claiming, index updates, and retry epochs

- touchLedgerIndex verifies its write and retries so a concurrent
  filer's merge cannot silently drop this key from the index.
- Foreign-ledger claims abort on any unreadable source instead of
  leaving it stranded once the new scoped key exists.
- A source that republished during the claim is left in place and its
  newer ids merge into the claimant's pending set.
- AddOrUpdate no longer overwrites the ledger epoch of an in-flight
  retry item, so its expiry or permanent outcome cannot forget a record
  that was re-queued after the attempt began.
- The recovery grace wait exits on shutdown instead of re-queueing
  after the filer has stopped.

* filer: requeue surviving records, persist claim deltas, guard index writes

- A dropped retry item (expired or permanent) whose ledger record was
  re-enqueued now pushes the id back through the hot queue instead of
  leaving it pending with nothing scheduled.
- Ids merged from a claim source that republished mid-claim are
  rewritten under our ledger immediately, so they are durable even if
  the claimant crashes before the next snapshot.
- touchLedgerIndex aborts when the index read fails for a real error;
  only ErrKvNotFound means the index is empty, so a transient failure
  can no longer wipe peer entries with a one-key write.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:37:31 +08:00
8d97284d0a filer: option to store system metadata logs in their own collection (#11551)
* feat(filer): option to store system metadata logs in their own collection

The filer's internal /topics/.system/log chunks are assigned to the
filer's default collection (-collection). In a multi-filer deployment
that default is often empty, so every restart flap, full-sync, or
event-buffered flush grows the default collection with system chunks that
are indistinguishable from user data in collection.list. This is a large
part of what makes the default collection balloon and confuses orphan
analysis.

This keeps the internal log in a dedicated collection when the operator
asks for one, without changing where user data goes:

- New optional override, filer.options.metaLog.collection (and
  .replication), read in NewFiler so both `weed filer` and
  `weed server -filer` honour it. Default "" => exactly today's
  behaviour (log follows the filer default), fully backward compatible.
- Resolution is a small helper: override first, then the filer default,
  then a storage rule matched on the log path. Kept separate from the
  user write path so the internal log targets itself.
- bucketCollection() is hardened the same way it already protects the
  filer's default collection: a bucket that happens to resolve to the
  redirected meta-log collection must not drop it on delete, because it
  backs internal log volumes.
- Scaffold filer.toml documents the new knobs under [filer.options].

Related to the persisted deletion ledger branch (fix/persist-deletion-queue):
together they cut the two sources of post-flap junk in the default
collection — that PR stops orphaned user-chunk leak on filer crash,
this one stops the internal log from living in default at all. They are
independent: no file overlap, no functional dependency; either can merge
first. They are paired only in the narrative of cleaning up default.

Adds unit tests for the collection/replication resolution chain, the
viper keys, and the bucket-delete guard (run green under -race).

Co-Authored-By: Athena 🏛️ <hermes-agent@local> (custom / Qwen3.8-Flash-Next-ROCmFP4)

* filer: collect bucket chunks when its collection survives the delete

bucketCollection returning "" preserves the collection, but the bucket
path still skipped per-entry chunk collection and could skip listing the
children entirely, so a bucket sharing the meta-log (or any preserved)
collection left its object chunks orphaned with no entry pointing at
them. Only the wholesale drop of a deleted collection skips those now.

Note in filer.toml that the meta-log target should stay stable: chunks
written under an older collection are not migrated.

* filer: exercise the metaLog override wiring through NewFiler

The viper test only echoed back the keys it set, so a wrong key in
NewFiler would still pass. It now asserts the fields NewFiler fills
from those keys.

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: tighten comments around the metaLog collection override

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 08:50:24 +08:00
Chris LuandDevin 90f8c4378f s3: restrict admin gRPC to local callers when no signing key (#11530)
* s3: restrict admin gRPC to local callers when no signing key

The S3 gateway's gRPC port (default 0.0.0.0:19000, always on) serves the
IAM cache and internal lifecycle admin services. checkAdminAuth was a
no-op when jwt.filer_signing.key was unset, so any reachable host could
PutIdentity an admin identity and take over the bucket data.

Without a shared key callers cannot be distinguished, so admin RPCs are
now limited to unix-socket, loopback, and the server's own interface
addresses. Remote filer-to-S3 propagation and lifecycle workers must set
jwt.filer_signing.key; the Bearer-token path is unchanged.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: fail closed on nil guard and refresh local addresses per call

Review feedback: a nil filerGuard bypassed all checks — treat it like a
missing key and require a local peer. The own-address set was cached
forever, so interfaces added later were rejected; enumerate per call
instead since admin RPCs are rare. Nil ctx is denied rather than panics.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: read the signing key once and bound interface enumeration

Review feedback: reading SigningKey twice could straddle a SIGHUP reload
— an old nonempty key skipped the local-peer check while the new empty
key verified the token. And enumerating interfaces per no-key call is
wasteful for co-located workers dialing the announced address; cache the
address set for 30s so new interfaces still become usable promptly.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: enumerate interface addresses per no-key admin call

A cached address set keeps trusting an IP after it is removed from the
host and reassigned to another machine — that host would then hold
unauthenticated admin access for the cache TTL. Per-call enumeration only
runs for non-loopback TCP peers on the no-key path, which is low-volume
admin traffic, so the freshness is worth the syscall.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 23:12:23 +08:00
Eliah RusinandClaude Opus 5.5 2a42d56437 ecbalancer: honour total-shards-per-rack cap in Place / PlaceDurabilityFirst (#11553)
* ecbalancer: honour total-shards-per-rack cap in Place / PlaceDurabilityFirst

Worker auto-EC encode places via Topology.Place, which capped each shard
type independently (ceil(data/racks), ceil(parity/racks)). On an 8-rack
topology that permits 3 total shards on one rack, so losing two racks
strands 6/14 and a 10+4 volume becomes unreadable.

- tryPlace caps the total shards (data + parity) per rack in both modes,
  whether or not ReplicaPlacement is set.
- rackTotalCap picks the smallest per-rack total the racks' real room
  (free slots, bounded by the per-disk cap and node free slots, counting
  shards already placed) can satisfy. On a uniform cluster it is
  ceil(shards/racks); a nearly full rack raises it just enough that the
  cap alone never fails an encode.
- PlaceDurabilityFirst gets a last rung that drops the rack cap
  ("rack-total-cap" in Relaxed), so it fails only when no disk has room.
  PlaceStrict keeps the cap as a hard limit.
- chooseShardDest tries the next rack when the chosen one has no node
  that fits, and room checks count the per-disk cap, so a rack whose
  disks are all at the cap is no longer picked and then failed on
  (pre-existing: 3-node rack + single-disk rack failed at shard 9).
- Docs no longer claim the cap guarantees surviving rack loss; the
  placement error names the caps in effect; the encode warning no longer
  says replica placement when other constraints were relaxed.

place_rack_cap_test.go covers 10+4 over 8 racks (max 2/rack, 3/rack on
master), a starved rack, nearly full racks, the preferred-tag tier, the
full-disk rack, and rackTotalCap directly.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ecbalancer: size the rack total cap from room left under SameRackCount

The rack total cap counted each rack's free disk room, but attempts that
enforce ReplicaPlacement also stop a node at SameRackCount shards. With
SameRackCount=1, four one-node racks and four three-node racks got cap 2,
which fits only 12 of 14 shards: strict placement failed and
durability-first relaxed replica placement although 1 per small rack and
up to 3 per large rack fits.

Attempts that enforce ReplicaPlacement now use a cap sized from each
node's remaining SameRackCount allowance; attempts that relax it keep the
disk-room cap.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 22:22:34 +08:00
Eliah RusinandClaude Opus 5.5 164c3db606 s3: return 403, not 500, when an over-quota bucket refuses a write (#11552)
* s3: return 403, not 500, when an over-quota bucket refuses a write

Filer AssignVolume flattened ErrReadOnly into the free-text
AssignVolumeResponse.Error string, so S3 PutObject / PutObjectPart via
UploadReaderInChunks could not match it with errors.Is and fell through
to 500 InternalError: retryable, and it hides the quota.

Add FilerError READ_ONLY and AssignVolumeResponse.error_code, set it
alongside the unchanged error text, and rebuild the sentinel with
filer_pb.AssignVolumeResponseError. weed_server.ErrReadOnly now aliases
filer_pb.ErrReadOnly so errors.Is matches on both sides, and
mapChunkedUploadErrorToS3Error maps it to ErrAccessDenied. There is no
"read only" substring matching, so a volume server's "volume N is read
only" stays retryable.

Carrying the verdict as a response code rather than a gRPC status keeps
clients from treating it as a transport failure: the S3 gateway does not
fail over across filers and the Java client does not retry it.

Wrap per-chunk copy errors with %w so CopyObject keeps the sentinel, and
map UploadPartCopy chunk errors through mapCopyErrorToS3Error instead of
always returning 500.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: re-run integration tests (PyPI download timeout)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 22:21:03 +08:00
Javier Garcia 8fdcf69eb0 s3api: report the stored checksum in GetObjectAttributes (#11529)
GetObjectAttributes accepted the Checksum attribute but never filled it
in, as its comment said SeaweedFS did not store S3 checksums. PutObject
and CompleteMultipartUpload store them now, and HeadObject returns them.
Fill in Checksum from the same entry fields, with the ChecksumType and
ChecksumCRC64NVME members the response did not have.

Also run ceph/s3-tests' test_get_checksum_object_attributes in CI.
2026-10-01 01:11:19 +08:00
Javier Garcia 988fc4f7ba s3api: do not store aws-chunked in an object's Content-Encoding (#11528)
* s3api: do not store aws-chunked in an object's Content-Encoding

aws-chunked in Content-Encoding names the SigV4 streaming framing of
the request body, which the gateway decodes on upload. PutObject and
CreateMultipartUpload stored the header as sent, so an object uploaded
with "gzip, aws-chunked" was served with that Content-Encoding, and one
uploaded with "aws-chunked" alone was served as aws-chunked. S3 drops
aws-chunked and keeps the other encodings.

Also run ceph/s3-tests' test_object_content_encoding_aws_chunked in CI.

* s3api: read every Content-Encoding field, and drop aws-chunked on copy

A client can send aws-chunked and the object's own encoding as separate
Content-Encoding fields. Only the first was read, so "aws-chunked"
followed by "gzip" left the object without its gzip. Combine all the
fields before dropping aws-chunked. CopyObject with the REPLACE
directive stored the requested Content-Encoding as sent: drop
aws-chunked there too.
2026-10-01 01:10:38 +08:00
ihnokim 11e8c4c288 master: follow heartbeat read-only changes in the layout's replica flag (#11527)
A replica's read-only flag in the volume layout only moved on registration
and on volume.mark. A change that arrived in the regular heartbeat updated
the node's record, which the writable list follows, but not the layout
flag, which the vacuum sweep reads. So the sweep kept trying volumes on a
disk that had gone read-only while the server ran, and after a restart it
skipped volumes that had since become writable again until the next
restart (issue #11516).

Apply the reported state to the flag for every changed volume. Only the
flag: the writable list stays with EnsureCorrectWritables and its
capacity guards.
2026-10-01 01:09:09 +08:00
ihnokim d3cd061c22 shell: say which read-only volumes volume.vacuum leaves alone (#11525)
* shell: say which read-only volumes volume.vacuum leaves alone

volume.vacuum without -volumeId runs the same sweep as the automatic
vacuum, which skips read-only volumes, and the master's response carries
no result. An operator whose disk filled up runs the command, sees it
return, and watches nothing change (issue #11516).

Before issuing the request, list the read-only volumes whose garbage is at
or above the threshold and point at -volumeId, which is the explicit path
PR #9861 opened for them. The help text says the same.

* shell: volume.vacuum hint survives a failed listing and looks at every replica

Review follow-ups: a failed topology listing no longer stops a sweep
without -volumeId, it only drops the hint; a volume counts as read-only
when any replica is, with the garbage ratio taken from the replica that
reports the most, which is what the sweep itself does; a converted index
that reports deletes without sizes is listed rather than hidden; and the
threshold is printed as given instead of rounded to two decimals.

* shell: do not guess a garbage ratio for a converted index

The master cannot compute one for a volume that reports deletes without
their sizes, and a guess of 1 would send the operator to -volumeId for a
volume the server may decline at that threshold. Leave it out and say so.
2026-10-01 01:08:48 +08:00
Khris RichardsonandClaude Opus 5.5 895d49b55b s3/iam: manage roles through the IAM API, with an opt-in persistent role store (#11522)
* s3/iam: manage roles through the IAM API, with an opt-in persistent role store

Roles could only come from the IAM config file: the S3 server pinned the
role store to memory and the embedded IAM API had no role actions, so a
role could not be created, retrusted or revoked without editing the file
and restarting every gateway.

Role store
- Read the `roleStore` key (the IAMConfig field already existed). With an
  IAM config file the default stays memory; with none it is the filer, as
  for OIDC providers, so zero-config clusters keep runtime-created roles.
- Roles from the IAM config file never go into a persistent role store,
  which outlives the file and may be shared by S3 servers with different
  files. They are served from memory beneath the store, as OIDC providers
  are: a stored role of the same name takes precedence, and deleting it
  restores the file's. A config-file role cannot be changed or deleted
  through the API (UnmodifiableEntity), and removing one from the file
  removes it at the next start. An in-memory store holds them as records,
  as before. They have no creation time, so CreateDate is omitted rather
  than reporting when this server started. SetRoleStore installs a store
  the same way, so a store set after startup keeps the config-file roles,
  as SetOIDCProviderStore does for providers.
- Watch /etc/iam/roles and drop the cached role definitions on change. The
  cached filer store otherwise serves a peer's stale role for up to its 5m
  TTL, which keeps a revoked trust policy in force on the other gateways.
- Role stores wrap ErrRoleNotFound for a missing role; the filer store
  used to report any failed lookup as "role not found". CreateRole proceeds
  only on a confirmed absence, so an unreadable store cannot let it write
  over an existing role.

IAM actions
- CreateRole, GetRole, ListRoles, DeleteRole, UpdateAssumeRolePolicy,
  AttachRolePolicy, DetachRolePolicy, ListAttachedRolePolicies. The reads
  are allowed in read-only mode.
- A role defined in the config file is reloaded from it at every start, so
  changing or deleting it through the API is refused (UnmodifiableEntity)
  rather than silently reverted.
- DeleteRole with policies attached is refused (DeleteConflict), as on AWS.
- Role names follow AWS's rules ([\w+=,.@-]{1,64}); a role is stored as
  <name>.json in the filer, so this also keeps a name from leaving the role
  store's directory. At most 10 managed policies per role (AWS's default
  quota; MaxManagedPoliciesPerUser is 10 too), LimitExceeded beyond.
- DeletePolicy is refused (DeleteConflict) while a role attaches the
  policy, as it already is for users and groups: roles attach policies by
  name, so a policy created later under the deleted one's name would
  otherwise take effect on the role.
- Role paths other than "/" and role tags are not stored, so they are
  refused rather than dropped.

Role IDs and sessions
- Roles get a unique RoleId when first stored (random, AWS AROA form),
  kept across updates; a config-file role gets a stable ID derived from its
  name, since it is created again at every start.
- Sessions issued through AssumeRoleWithWebIdentity, AssumeRoleWithCredentials
  and AssumeRole carry the role's ID (claim "rid"), and a request under a role
  whose current ID differs is denied. Resolving a session's policies by role
  name let a session outlive its role: once a role was deleted, a role later
  created under the same name — with a different trust policy and different
  policies — revived every unexpired session of the old one with the new
  role's permissions. Sessions issued before this change carry no ID and are
  unaffected until they expire.

Integration test (test/s3/iam, run with `make start-services`):
TestWebIdentityWithProviderAndRoleManagedThroughIAMAPI configures an OIDC
provider, a managed policy and a role entirely through the IAM API against a
JWKS served by the test, then checks the trusted subject gets credentials
scoped to the attached policy; another subject, a token signed by another
key, an unsigned token and a token for another audience are refused; and UpdateAssumeRolePolicy moves the
trust at once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: bind every role session to its role, and change roles atomically

Review follow-ups.

Session binding
- The role-ID check ran only when a session carried no policy names, and
  AssumeRole embeds the role's attached policies, so those sessions kept
  their permissions after the role was deleted or recreated. The check
  now runs for every session carrying a role ID, before policy selection.
- A named role that cannot be resolved at issuance gets no session,
  instead of one with no role ID (which nothing binds).
- A config-file role's ID is derived from its name and trust policy, not
  the name alone: a different role put in the file under the same name
  gets a new ID, while an unchanged role keeps its sessions across restarts.

Role writes
- RoleStore gains UpdateRole, a read-modify-write that lands only if the
  role is unchanged since the read, and otherwise re-reads and retries. The
  filer store uses the filer's write conditions (IF_NOT_EXISTS for a new
  role, IF_ENTRY_EQUAL otherwise). CreateRole, UpdateAssumeRolePolicy and
  Attach/DetachRolePolicy all go through it, so two gateways no longer
  overwrite each other's changes, a change racing a delete no longer
  writes the role back, and of two concurrent creates one gets
  EntityAlreadyExists.
- The filer store's ListRoles pages past 1,000 entries and fails on a
  broken stream instead of returning what arrived, so DeletePolicy's
  attachment check sees every role. ListRoles skips a role deleted between
  listing and reading it.
- CreateRole validates first; a failed write is ServiceFailure, not
  InvalidInput. Any Tags.* parameter is refused, not only the first key.
- ExecuteAction's skipPersist covers the S3ApiConfiguration only; the
  comment now says so. Role and OIDC provider actions write their own stores.

Each fix has a test that fails without it. Against a real filer with two
gateways, concurrent AttachRolePolicy calls lost 1-4 of 8 attachments per
run before this change and none after.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: one role snapshot per decision; DeleteRole is atomic; watch a custom role store path

Review follow-ups.

- Authorization evaluates the policies of the role definition the session's
  binding was checked against, instead of reading the role again: a role
  replaced in between cannot lend a session its policies.
- AssumeRole and AssumeRoleWithLDAPIdentity issue the session from the
  definition whose trust admits the caller (IAMManager.ResolveRoleForPrincipal),
  and take its ID, duration cap and embedded policies from that same
  definition. A role replaced after the caller's trust check by one that does
  not trust the caller now yields AccessDenied, not a session bound to the
  replacement.
- A RoleUpdate that returns nil deletes the role, on the same condition as a
  write: the filer store deletes with ObjectTransaction on IF_ENTRY_EQUAL,
  routed and locked like the conditional CreateEntry. DeleteRole decides
  against the role it deletes, so a policy attached meanwhile on another
  server is a DeleteConflict, and a delete never removes a role written
  after its check.
- S3 servers watch the role store's configured basePath, not only
  /etc/iam/roles, so a custom path also drops peers' cached roles on change.

Each has a test that fails without it. Live against a real filer: DeleteRole
refuses while a policy is attached and removes the entry once detached; all
test/s3/iam CI stages pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: state which roles DeletePolicy's attachment check can see

RolesAttachingPolicy sees the stored roles and this server's config-file
roles. A role defined only in another server's IAM config file is invisible
to it, so a config-file role that attaches a managed policy is protected
only on the servers whose file defines it. The doc comment now says so and
how to avoid it: keep such roles in every server's file, or attach only
config-file policies to config-file roles.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: note that a role store set after startup is not watched for peer changes

S3 servers build their metadata watch list once, at startup, from the role
store installed then. SetRoleStore's doc now says that a filer-backed store
installed later with a different basePath is not watched, so peers' changes
to it reach this server's cached roles only when the cache expires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:45:47 +08:00
Khris Richardson 95e0b74fb6 s3/iam: retry a failed OIDC provider refresh until the store answers (#11521)
* s3/iam: retry a failed OIDC provider refresh until the store answers

RefreshOIDCProvidersFromStore reports a failure and nothing retries it. Its
callers can't: a metadata-subscription event reports each change once, so a
refresh that found the filer unreachable on it (the filer restarting, say)
left a peer's new provider untrusted, or a deleted one trusted, until some
unrelated later change. The refresh after a local IAM API mutation has the
same shape. Only the startup load retried.

A failed refresh now retries in the background with the startup load's
backoff until the store answers. At most one retry runs, however many
refreshes fail meanwhile, and installing another store cancels it. The
startup load uses the same path instead of its own.

Seen on a SeaweedFS operator cluster whose filer restarted while an
S3OIDCProvider was created: the gateway logged "OIDC provider refresh after
/etc/iam/oidc-providers change failed: ... fail to dial". The operator's
periodic re-apply happened to recover it; an IAM API client would not.

* s3/iam: never retry or apply a superseded OIDC provider store, and never drop a failure during a retry

Review of the retry (#11521) found two ways to lose the state it protects.

A refresh of store A that failed as store B was installed could start a
retry for A after B's install had cancelled retries. Nothing cancelled it,
and when A answered it replaced B's providers in STS. The installed store
now changes under the retry lock, a store that is no longer current gets no
retry, and a snapshot of a replaced store is never handed to STS, even when
the refresh listed it just before the swap.

A refresh that failed while a retry ran was dropped by the at-most-one
guard, though the retry might already have listed an older snapshot, so the
change the failed refresh would have loaded stayed unloaded. The retry now
runs once more after its success when a failure arrived meanwhile.

Each has a test that fails without its guard.
2026-09-30 17:33:40 +08:00
Chris LuandDevin 0978e7f833 vacuum: keep disk-full read-only volumes reclaimable (#11519)
* storage/topology: keep disk-full read-only volumes vacuumable

The vacuum sweep skipped every read-only replica, so a volume that went
read-only because its disk filled could never reclaim its garbage — the
exact situation compaction exists for. The volume server now reports
disk_space_low in VacuumVolumeCheckResponse, and the sweep skips a
read-only replica only when the flag is clear. An explicit volumeId
vacuum is unaffected: it already bypassed the read-only rule.

The field takes number 4: 2 and 3 are downstream-allocated for tombstone
retention, keeping the wire merge clean.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* storage: measure vacuum free space against live bytes

The pre-compaction space check required the current .dat + .idx size
free, which includes the garbage being reclaimed — on a nearly full disk
that estimate can never fit, so the volume stayed garbage-bound forever.
Measure against the estimated compacted output instead: superblock plus
live index entries plus live content bytes, with the existing ten
percent buffer unchanged. Mirrors the same check in the Rust volume
server.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: count per-needle framing in the compacted-size estimate

The live-bytes estimate covered each live needle's content and index
entry but not its .dat framing (header, checksum, timestamp, padding —
~32 bytes on version 3). For small-needle volumes that is more than the
10% headroom, so a disk with space between the estimate and the real
output still ran out mid-compaction. Rust side mirrors the same formula.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* storage: report disk_space_low only when it is the sole read-only cause

Review feedback (ihnokim, greptile, devin): a volume read-only for low
disk space AND an operator mark or I/O quarantine was still eligible for
the automatic sweep, rewriting a copy meant to stay protected. The flag
now reports only the benign sole-cause case in both servers.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* topology: fail closed when the read-only lookup misses in the sweep

A heartbeat can drop the volume from the DataNode cache between the
location-list copy and VacuumVolumeCheck; a lookup error previously
skipped the read-only check entirely. Review feedback (coderabbit).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 17:32:41 +08:00
Alex Hu 38ce95d960 s3api: always write XML timestamps with three fractional digits (#11520)
CopyObject responses carried LastModified values such as
"2026-09-29T20:30:04.56Z": trailing zeros of the fractional seconds were
trimmed, and a whole-second value had no fraction at all. AWS S3 always
writes exactly three digits ("...04.560Z"), and clients that parse with a
fixed-width pattern reject anything else. minio-java 8.6.0
(yyyy-MM-dd'T'HH:mm:ss.SSS'Z') throws DateTimeParseException, so roughly
one CopyObject in ten fails on the client even though the copy succeeded.

Two causes:

- xsdDateTime marshalled with "2006-01-02T15:04:05.999999999", which
  drops trailing zeros. It now writes UTC with ".000Z".
- CopyObjectResult.MarshalXML had a pointer receiver, but the handlers
  pass the result by value, so encoding/xml never called it and fell back
  to time.Time's RFC 3339 encoding. It now has a value receiver.
  CopyPartResult had no custom marshaller at all; it now uses xsdDateTime.

Follow-up to #8394 / #8398, which truncated these timestamps to
milliseconds but kept the trimmed format.
2026-09-30 16:22:31 +08:00
Chris LuandDevin 757917f564 filer: evict remote-cached objects under storage pressure (#11515)
* filer: identify remote-mounted entries safe to drop under disk pressure

ListEvictableRemoteEntries walks every mounted directory directly on the
filer store (no lazy remote listing) and returns entries that hold local
chunks fully synchronized with remote, ordered oldest-cached first.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: evict remote-cached chunks oldest-first and vacuum the garbage

uncacheRemoteEntry applies the same transition remote.uncache does -
cleared chunks plus a reset LastLocalSyncTsNs under the entry path lock -
and evictRemoteCachedEntries serializes passes over all mounts until a
byte target is met. Aged victims are preferred; a second pass accepts any
synchronized cached entry when aged ones cannot cover the request, since
a failed read is worse than a dropped hot object.

Cleared chunks only become disk space after compaction, so
reclaimRemoteCacheSpace pairs each pass with a rate-limited VacuumVolume
call that also picks up orphaned partial fills.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: trigger remote cache eviction under storage pressure

A periodic check (30s) reads disk usage from master topology and evicts
remote-mounted cached chunks once any disk crosses
-filer.remoteCacheEvictThreshold (default 0.9; 0 disables), with a vacuum
pass to reclaim the tombstoned needles.

The cold-read cache path also kicks the same reclaim when a fill fails on
exhausted volumes - the request still falls back to streaming from the
remote, but the cache stops being permanently wedged full.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: flush deletion queue before remote cache vacuum

Vacuum ran immediately after eviction while evicted file IDs still sat
in the asynchronous deletion queue, so compaction saw no garbage and the
cache stayed wedged. Flush the queue synchronously first and shorten the
vacuum cooldown so sustained pressure does not wait five minutes between
reclaim passes.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: cover remote cache eviction under capacity pressure

Unit tests pin the eligibility filter and oldest-first ordering; the
integration test runs a constrained two-node setup that saturates the
cache, verifies the oldest synced entry is evicted and vacuumed, and
that a later read re-caches it.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: coalesce remote cache reclaim passes

A failed cache fill used to queue behind any in-flight eviction,
stacking full mount traversals during a write-failure storm. Skip the
pass when one is already running; the caller falls back to streaming
from remote regardless.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: stop the remote cache janitor on shutdown

The eviction ticker kept running after Shutdown closed the metadata
store and could traverse a closed store. Give the janitor a context
cancelled from Shutdown and propagate it into its master RPCs and
traversals.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: vacuum only tombstoned volumes and retry deferred passes

VacuumVolume with no volume id swept every collection, compacting
volumes unrelated to the cache fill that failed. Now the reclaim path
collects the vids of file ids actually flushed from the deletion queue
and compacts only those. Vids that land inside the vacuum cooldown stay
in a pending set the janitor retries on each tick, so chunks evicted
just after a sweep are not stranded until the next pressure event.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: count only pressured disks when evicting remote cache

The janitor measured the largest excess on one disk but let bytes on
healthy disks satisfy the reclaim target. Split the topology disk view
per physical disk and count only chunk bytes whose volumes sit on an
over-threshold disk; entries contributing nothing there are skipped.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: compare remote cache sync time at nanosecond precision

Second-precision mtime comparisons let a local write in the same second
as the last sync still qualify as evictable, discarding unsynced
changes. Compare LastLocalSyncTsNs against full-precision mtime
(mtime_ns round-trips through the entry codec), and apply the same fix
to remote.uncache's inline check.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: invalidate remote sync stamp on local content change

A local overwrite that keeps the remote entry's LastLocalSyncTsNs looks
evictable even though the remote copy no longer matches, and some write
paths stamp mtime at second precision so a timestamp comparison cannot
catch it. UpdateEntry now clears the stamp when chunks change without a
fresh stamp, leaving replicated updates authoritative.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound remote cache master rpcs and vacuum all evicted garbage

VolumeList and VacuumVolume now run under a 30s context so a stalled
master cannot wedge the eviction janitor. The targeted vacuum drops the
garbage threshold so volumes with under 10% deleted bytes still compact.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: tolerate straggler fills in remote cache eviction test

Detached fills from the concurrent wave keep racing the final checks:
live chunks legitimately fill both volumes, and a re-cached object can
be evicted again before its commit is observed.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: start remote cache eviction loop after filer init

The janitor's first tick dereferences fs.filer; starting the goroutine
before NewFiler assigns it could panic when startup exceeds an interval.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: keep remote cache vacuum intent across retries

Evicted entries now record their chunk volumes for vacuum directly, so
the intent survives whoever consumes the shared deletion queue first.
A pending volume keeps several vacuum attempts so tombstones that land
late are still compacted, and the janitor retries pending volumes under
the reclaim mutex instead of flushing unrelated deletes every tick.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound each remote cache vacuum request independently

A shared 30s deadline across pending volumes let one slow compaction
cancel the rest. Each VacuumVolume now gets its own context, and pending
volumes keep more attempts since the master reports request acceptance
rather than compaction.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: tighten remote cache reclamation bound

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: treat chunk timestamp changes as content changes

chunksEqual now also compares ModifiedTsNs so an update that rewrites a
chunk record still invalidates the remote sync stamp.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: run remote cache queue flush under the reclaim context

BatchDelete for flushed file ids now uses the caller's context instead of
context.Background(), so a reclaim pass bounded by shutdown or timeout
stops its deletes too. Other callers keep their existing behavior.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: scope remote cache vacuum to evicted volumes

The flush no longer feeds the shared deletion queue's ids into the
pending set — only evicted chunks' volumes are tracked, so ordinary
deletions no longer pick up repeated vacuum attempts. The flush also
runs under a shutdown-immune bounded context and is skipped when no
volume is pending.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: retry remote cache vacuums even after unmount

Pending volumes were only retried while a remote mount existed; removing
the last mount skipped every later pass and left evicted bytes allocated.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: reclaim partial cache fills that run out of capacity

A fill that fails midway queues its written chunks for deletion, but
when no entries remain evictable the reclaim pass found no pending
volumes and skipped the flush and vacuum entirely, leaving the partial
garbage to the slow periodic vacuum while the disk stayed full. Mark
the failed fill's chunk volumes pending so the pass tombstones and
compacts them even when nothing was evicted.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 22:05:42 +08:00
5da137233d s3/iam: persist IAM-managed OIDC providers in the filer, and trust them after a restart (#11510)
* s3/iam: persist IAM-managed OIDC providers in the filer, and trust them after a restart

The S3 server's IAM config loader never read the documented
`oidcProviderStore` key, so the OIDC provider store was always in memory:
a provider created with CreateOpenIDConnectProvider lived in one gateway's
process, was lost on restart, and was never seen by peers. The
/etc/iam/oidc-providers metadata subscription refreshed from that empty
in-memory store.

- Read `oidcProviderStore` and pass it to the IAM manager. With an IAM
  config file the default stays memory. With no config file (zero-config
  IAM, as `weed filer -s3` and operator-managed clusters run) it defaults
  to the filer: there is nothing static to shadow, and providers created at
  runtime otherwise vanish on restart.
- With a store that outlives the process, load the STS runtime view from it
  at startup, so providers created on an earlier boot or on a peer are
  trusted without waiting for the next mutation.
- If the store cannot be read at startup (a filer not up yet), the load is
  retried in the background with backoff until it succeeds: the metadata
  subscription reports only later changes, so providers already stored would
  otherwise stay unknown to STS until one of them changed.
- Mark records mirrored from STS.Providers as `source: static-config`, and
  at startup delete such records whose provider has left the config, so
  removing a provider from the config file still revokes it. Records created
  through the IAM API are never pruned.
- The filer store reported every failed lookup, an unreachable filer
  included, as ErrOIDCProviderNotFound, which CreateOIDCProvider reads as
  "free to create". Only a confirmed absence is now not-found.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: keep config-file OIDC providers out of a persistent store

Review of the previous commit found that mirroring the IAM config file's
providers into a persistent store, and pruning them when they leave the
file, breaks as soon as S3 servers share a filer:

- a server prunes stored config-file providers its own file does not list,
  including ones a peer's file still defines (a zero-config server prunes
  them all);
- mirroring overwrites an API-created provider with the same ARN and marks
  it config-owned, so a later prune deletes it;
- a failed mirror write or a failed prune leaves a stale record trusted;
- a mirrored record is loaded into STS at startup as an IAM-managed provider
  and shadows the config-file provider, dropping the settings a record does
  not carry (jwksUri, roleMapping, policyClaim, ...).

A persistent store now never receives the config file's providers. STS keeps
serving them from its static configuration, as it always has; the IAM API
lists and returns them from memory, refuses to change or delete them
(UnmodifiableEntity; change them in the file) and to create another provider
with their ARN (EntityAlreadyExists). The store holds only providers created
through the IAM API, and those are what startup loads into STS. There is
nothing to prune, so the source marker is gone. An in-memory store keeps its
behaviour: the config file's providers are records in it, as before.

buildOIDCProviderFromRecord also carries PolicyClaim and
AllowedPrincipalTagKeys now; they were dropped whenever an API-created
provider was loaded into STS.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: send UnmodifiableEntity as a 400, not an internal error

The IAM API's error writer had no case for UnmodifiableEntity, which the
previous commit returns for a change to a config-file provider, so it went
out as a 500 ServiceFailure that clients retry. AWS sends it as a 400.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: document stored-over-config precedence, drop invented CreateDate, cancel superseded retries

Follow-ups from review of b881982d2:

- A provider stored under the same ARN as a config-file provider takes
  precedence in the IAM API, matching STS, which already prefers
  IAM-managed providers so that an API call can shadow a bootstrap entry.
  Deleting the stored provider brings the config-file one back. This was
  already the behaviour; it is now documented and tested.
- A config-file provider no longer reports its server's start time as
  CreateDate, which changed on every restart; GetOpenIDConnectProvider now
  omits the date for it. An in-memory store still stamps its copies at load,
  as before.
- The startup retry runs under a cancellable context, is cancelled when
  another store is installed, and retries the store it was started for
  rather than reading the manager's field, so replacing the store neither
  leaves the old retry running nor races with it (go test -race).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: serialize OIDC provider refreshes so an older snapshot cannot restore a deleted provider

Refreshes run concurrently: after an IAM API change, on a peer's change
and in the startup retry. Each lists the store and then hands STS the
result, so a refresh that listed before a DeleteOIDCProvider could finish
after that call's own refresh and keep the deleted provider trusted until
the next change. Refreshes now hold a lock from the read to the hand-off,
and a startup retry cancelled by installing another store drops its
snapshot instead of applying it.

The retry-cancellation test waits for the retry by polling instead of a
fixed sleep.

* s3/iam: route SetOIDCProviderStore through installOIDCProviderStore

A store installed after Initialize skipped the static-provider overlay
and startup hydration: config-file providers disappeared from the IAM
API, ErrOIDCProviderStatic no longer protected them, and stored
providers were never trusted until the next mutation or peer event.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:51:39 +08:00
Chris LuandDevin 62d4f9152a iam: evaluate trust policies deny-by-default (#11513)
* iam: evaluate trust policies deny-by-default

EvaluateTrustPolicy seeded its result with the engine's DefaultEffect,
so a non-matching trust-policy statement set still resolved to Allow
when the IAM config sets policy.defaultEffect=Allow. A caller holding
a validly signed token from a registered provider could then assume a
role its trust policy does not admit.

Trust policies now start from implicit deny, matching AWS semantics and
the pre-d751623 behavior of evaluateTrustPolicy; DefaultEffect still
governs identity-policy evaluation.

Upgrade note: deployments on defaultEffect=Allow whose trust policies do
not match their callers will see those assumptions refused.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: cover trust policy implicit deny under DefaultEffect=Allow

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:36:01 +08:00
Chris LuandDevin a901c1a5e2 filer: compare IF_ENTRY_EQUAL chunks by fid, not file_id (#11514)
The stored entry came through FindEntry, which restores chunk file ids
from their fid form, while an expected entry built from a metadata-log
event still carries the serialized form (file_id moved into fid). The
proto.Equal saw file_id "" against the restored id and refused every
stamp, so remote.sync re-uploaded each entry and the RemoteEntry stamp
never landed.

Clone both sides and run BeforeEntrySerialization before comparing, so
chunks match on their fid and the file_id spelling is ignored; the stored
entry and the request's ExpectedEntry are left untouched.

Generated with [Devin](https://devin.ai)

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:34:07 +08:00
github-actions[bot] 530be3e373 4.48 2026-09-28 15:53:40 +00:00
Chris LuandDevin 9b3b12c607 filer: pin-aware reader cache eviction and stream release (#11503)
* filer: synchronize stream pins and release them on transitions

Guard chunkStream.cacher with the ReaderCache lock everywhere: mount
sections share one ChunkReadAt across concurrent reads, and unsynchronized
release could double-unpin. Reads served from the chunk cache now detach
the stream's pin instead of retaining the previous chunk. Eviction prefers
unpinned downloaders so a pinned buffer is not dropped mid-stream. A new
ReleaseStream lets callers drop their pin without destroying the shared
cache; S3 and WebDAV readers use it. lastChunkFid becomes atomic since
concurrent mount reads can update it.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: keep eviction bounded when every downloader is pinned

Both eviction paths still fall back to a pinned victim when no unpinned
one exists, so abandoned stream pins cannot bypass the downloader limit
or stall the memory budget. Budget eviction also rechecks the pin under
the ReaderCache lock at removal time: a stream that pinned the selected
victim in between keeps it mapped and the selection retries.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: restore budget bookkeeping when a victim gets pinned mid-eviction

removeUnpinned losing the pin race left the victim out of the idle list
while still holding its reservation, making it unevictable even as the
pinned fallback. Push it back when the reservation is still live.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 22:37:04 +08:00
Tobias Gurtzick 3abdef3202 filer: keep shared chunk buffers pinned while another stream reads them (#11502)
The ReaderCache is shared by all streams of a process (every S3 GET, for
instance), but a ChunkReadAt released chunks as if it owned them:

- moving on to the next chunk called UnCache on the previous one,
  destroying the buffer even when other streams were still inside it;
- since #11384 a buffer is dropped once any reader has consumed it to the
  end and no read call is in flight. Streams copy out in slices (256 KiB
  in the S3 gateway), so between two calls a slower stream is not
  attached and loses the buffer to a faster one.

Either way the slower stream refetches the whole chunk from the volume
servers. With many clients downloading the same popular object at once,
each chunk is fetched over and over; in production we saw the S3 gateway
pull ~10 Gbit/s from volume servers while serving ~1 Gbit/s to clients.

A ChunkReadAt now pins the chunk it is positioned in. The pin is taken
and released only under the ReaderCache lock, since concurrent ReadAt
calls on one ChunkReadAt (as in mount) share it. It is released when the
stream reads the chunk to its end, moves to another chunk (including one
served from the chunk cache), or falls back to random reads. A buffer is
dropped once no stream pins it and no read is in progress, if it was
consumed or its last stream left it; a read still in flight when the
stream leaves drops it on detach, as UnCache did via destroy. Eviction by
slot limit and memory budget is unchanged.

lastChunkFid is now guarded as well: concurrent ReadAt calls raced on it.

Tests: two ChunkReadAt instances streaming one object in interleaved
slices fetch each chunk exactly once (2-3 times before); leaving a chunk
for a chunk-cache hit or while another read is in flight releases it;
concurrent ReadAt calls on one ChunkReadAt leave no pins behind under
-race.
2026-09-28 21:56:04 +08:00
Chris LuandDevin 43fd5b8d82 volume: reclaim staged EC shard generations left by the 2PC switch (#11501)
* volume: remove staged EC generation files on teardown and shard delete

The 2PC generation switch stages each run as <base>.ecNN.v<N> plus
versioned .ecx/.ecj/.vif files. Nothing on the volume server removes
them: isEcDataShardFile only recognises the exact .ecNN name, so the
staged files are invisible to every bookkeeping pass, and even
full_teardown's wipe-all path left them behind. Each re-encode therefore
leaks a full shard set per shard-holding disk.

RemoveEcGenerationFiles sweeps <base>.ec*.v<N> and <base>.vif.v<N>,
optionally keeping generations at or above a threshold; teardown and the
reconcile wipe remove every generation, and a per-shard delete removes
that shard's staged generations too.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: delete staged EC generations older than N via VolumeEcShardsDelete

After a 2PC generation switch commits, the superseded generation's
<base>.*.v<N> files sit on disk with no cleanup path: teardown removes
everything, and a per-shard delete only touches the named shards, so the
executor had no RPC that reclaims just the staged leftovers.

delete_generations_older_than removes staged generation files strictly
below the threshold on every disk. Versioned files are never mounted, so
nothing is unloaded first; the committed generation and the canonical
files are preserved.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: mirror staged EC generation cleanup

Parity with the Go volume server: remove_ec_generation_files sweeps
<base>.ec*.v<N> and <base>.vif.v<N> staged by the 2PC switch, called by
remove_ec_volume_files (which covers both teardown paths) and the new
delete_generations_older_than request field; delete_ec_shards removes a
shard's staged generations along with the canonical file.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: match staged generation filenames literally

filepath.Glob interprets metacharacters in the collection part of the
base name, so a collection like a[bc] could match another volume's
staged files (or miss its own). Scan the directory and compare names
literally instead, mirroring the Rust read_dir implementation.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: report generation-sweep errors and drop the store lock first

- snapshot the location base names under the read lock and run the
  filesystem sweep after dropping it, so a slow disk cannot stall the
  store;
- record per-entry read_dir errors in remove_ec_generation_files and
  propagate them from remove_ec_shard_generations instead of flatten()
  skipping them;
- warn when a staged-shard generation fails to delete rather than
  reporting success with files left behind.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: fail shard delete when the staged-generation listing fails

A transient ReadDir failure fell back to removing canonical shard names
only: staged .v<N> files survived while the RPC still reported success,
leaving the leak invisible to retrying callers. ENOENT still means the
disk simply has no such directory; other listing errors now propagate.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: propagate staged-generation removal failures

delete_ec_shards logged remove_ec_shard_generations errors and the RPC
returned success while staged .v<N> files remained, diverging from the
Go handler which surfaces the failure. The sweep keeps processing the
remaining shards, retains the first error, and volume_ec_shards_delete
maps it to Status::internal so callers can retry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: notify state change even when the shard sweep errors

delete_ec_shards already deletes and unmounts the shards before
returning a staged-generation failure, so returning early skipped
volume_state_notify and the master kept routing to them until the next
heartbeat. Notify before propagating the error.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:55:25 +08:00
yi111andChris Lu 150a69fe11 master: make volume capacity reservation timeout configurable (#11426) (#11497)
* master: make volume capacity reservation timeout configurable (#11426)

* master: expire reservations on reads, fix int timeout units

- AvailableSpaceForReservation now expires reservations too: a node that
  is full of reservations is filtered out before TryReserveCapacity can
  clean them, which stranded expired capacity indefinitely.
- Drop TryReserveCapacityWithTimeout: a per-call timeout lets one caller
  expire another's live reservations, and the Node interface stays
  stable for implementations outside this tree.
- parseReservationTimeout no longer routes integer values through
  GetDuration, which read them as nanoseconds; bare numbers are
  seconds. The 5m fallback is now the shared DefaultReservationTimeout.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-28 18:40:11 +08:00
yi111andChris Lu 4fec65d949 filer: demote client-cancelled directory listing log from error (#11495) (#11496)
* filer: demote client-cancelled directory listing log from error (#11495)

* filer: quote path in canceled listing log

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-28 14:35:41 +08:00
Chris LuandDevin f9289f0570 s3: do not promote ?prefix into the object for non-List actions (#11494)
* s3: do not promote ?prefix into the object for non-List actions

authRequestWithAuthType mapped an empty object to the prefix parameter for
every action, so PUT /bucket?versioning&prefix=x authorized as Write:bucket/x.
An object-scoped grant (Write:bucket/*) could then change bucket versioning,
lifecycle, cors, and object-lock configuration, and the promoted object also
made ResolveS3Action report s3:PutObject to attached IAM policies.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: treat GET ?uploads as a bucket listing for authorization

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: resolve the listing action through the bucket-level object

resolveS3AuthTarget fed the promoted prefix to ResolveS3Action, so a
bucket-level ?uploads request resolved as s3:GetObject on the prefix ARN
in the admin explicit-deny check. Resolve both action and resource
against the object the bucket listing actually scopes.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: resolve the listing action through the bucket-level object in AuthorizeAction

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: drop the unreachable object-level uploads case from the resolver test

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 07:17:58 +08:00
4303b3aa4c s3: keep a listing's start position inside the requested prefix (#11493)
* s3: a list marker that sorts past the prefix leaves nothing to list

AWS scopes a listing to keys under Prefix; StartAfter, Marker and
continuation tokens only reposition inside that range. A marker that
diverges from the prefix at a larger byte is after every key the prefix
can match, so the page is empty. normalizePrefixMarker used to keep such
a marker as the walk cutoff at the bucket root, where the walk descends
into the marker's own directory and returns keys the prefix never names.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: keep the listing variant's action when a prefix is promoted to object

authRequestWithAuthType promotes ?prefix= into the object argument for
the legacy CanDo path. ResolveS3Action treats a non-empty object as
object-level, so a bucket-level ?versions or ?uploads request carrying a
prefix missed its specific action and fell back to the base List action:
an s3:ListBucket grant then covered s3:ListBucketVersions, and an
explicit Deny on the specific action was skipped on the same path.

Resolve the action against the same bucket-level object the resource
ARN already uses.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: treat GET ?uploads as a bucket listing for authorization

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Update weed/s3api/auth_credentials.go

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-28 06:58:44 +08:00
yi111andChris Lu a0ee7ba314 s3: ignore empty intermediate directories in bucketHasUserObjects (#11490) (#11491)
* s3: ignore empty intermediate directories in bucketHasUserObjects (#11490)

* s3: keep nested reserved-named dirs from hiding user objects

Reserved folders (.uploads, *.versions) are internal only at the bucket
root; deeper entries with those names are user key prefixes and must be
walked. Also treat a missing subdirectory as empty via isFilerNotFound
(list errors cross gRPC as status errors, not the sentinel), let names
containing backslashes count as objects, and walk iteratively so empty
chains deeper than the old scan depth no longer report non-empty.

* s3: treat reserved-named directories as internal at every level

Object listing interprets .uploads and *.versions directories as
internal storage wherever they appear, so walking them during the
emptiness check would report invisible version remnants as user objects
and block deletion. A reserved name on a file still counts, matching
listing which only special-cases directories.

* s3: count explicit directory objects under reserved names

A directory object created by PutObject (MIME or prefix-object marker
set) is user data even when named .uploads or *.versions; only a plain
directory with a reserved name is internal storage.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-28 06:56:46 +08:00
Chris LuandDevin a976b21010 s3: require dedicated object-lock permissions for x-amz-object-lock-* headers (#11492)
* s3: require dedicated object-lock permissions for x-amz-object-lock-* headers

PutObject, CreateMultipartUpload, and PostPolicy honor the retention and
legal-hold headers after only the route's s3:PutObject check, so a
write-only principal could pin a version under COMPLIANCE retention that
nobody can remove before its retain-until date. On AWS these headers
require s3:PutObjectRetention / s3:PutObjectLegalHold. validateObjectLockHeaders
is the shared funnel for all four call sites; it now authorizes the
corresponding dedicated action when each header is present.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: record the verified POST-policy signer as the request identity

The handler authenticated the form policy signature but stored only the
signer's name, so downstream authorization (the object-lock header check)
re-authenticated the form-signed request as anonymous and evaluated the
wrong principal.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 20:17:12 +08:00
Chris LuandDevin a261f90e18 vacuum: let the sweep release volumes that stay empty and quiet (#11477)
* vacuum: let the sweep release volumes that stay empty and quiet

Vacuuming reclaims bytes but not slots: a fully emptied volume stays
registered to its collection forever, and since growth is gated only on
slot count a store at 99% free disk can still refuse writes to other
collections (#11429). volume.deleteEmpty exists but is manual-only.

With -vacuumDeleteEmptyAfterSeconds (or master.vacuumDeleteEmptyAfterSeconds
under weed server/mini; default 0, off) the automatic sweep now deletes
replica copies that have stayed empty and quiet for that long, the same
rule volume.deleteEmpty applies on demand: remote-backed copies are
skipped, and every delete carries the volume server's onlyEmpty /
onlyGarbage guards so a copy written since the last report is refused
rather than removed. Copies that still hold data or were written
recently stay; only a volume whose every copy is deleted leaves the
sweep's work map, sparing a compaction of bytes that are all deleted.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: harden empty-volume sweep against partial and racing deletes

Review follow-up on #11477:

- delete a volume only when every replica copy is a verifiable
  empty-and-quiet candidate; deleting the empty copy of a volume whose
  sibling holds live files would silently cut its replica count
  (greptile P1).
- drain the volume out of the writable list before deleting, the same
  drain the compact pass uses, so PickForWrite stops assigning it and
  pending writes settle (devin).
- bound the VolumeDelete RPC so one stalled server cannot hold the
  vacuum lock indefinitely (greptile P1, reusing allocateVolumeTimeout).

The vid2location panic scenario raised in review does not exist:
VolumeLocationList methods are nil-receiver safe and a missing vid just
fails enoughCopies, so a partially deleted volume skips compaction
instead of crashing the sweep.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: unregister deleted empty replicas and prune the sweep list

A successful VolumeDelete only updates the volume server; the master
still tracked the replica and kept it in the sweep's location list for
the compaction pass (coderabbit on #11477). Unregister the replica right
after its delete succeeds and drop it from the sweep copy, so a partially
deleted volume only compacts copies that still exist.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: pin deleting volumes out of the writable list across heartbeats

Review follow-up on #11477 (greptile): DrainAndRemoveFromWritable only
removed the volume once; a heartbeat landing between the drain and the
replica deletes re-evaluated writability and re-added it, so a client
write could reach a replica whose siblings were already gone and leave
the volume under-replicated when the last copy refused its onlyEmpty
delete.

MarkDeleting records the vid in deletingVolumes — checked inside
setVolumeWritable so heartbeat, capacity-recovery, and admin re-add
paths all hold it out — and UnmarkDeleting releases it once the sweep
finishes the copy pass. A partially deleted volume's surviving replicas
then return to writable through the normal heartbeat path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: restore writability when a sweep delete survives

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 07:05:22 +08:00
Chris LuandDevin be29f44d87 s3: record requester identity before the authz verdict (#11479)
* s3: record requester identity before the authz verdict for audit

Identity was only stored in request context on the success branch, so
denied requests reached WriteErrorResponse without requester attribution
and audit entries had empty requester/requester_arn/requester_identity.
Authentication failures still resolve no identity, so unauthenticated
denials stay unattributed.

Fixes #11474

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: keep the resolved identity through authz denial in Auth

Review follow-up on #11479 (devin): authRequest discarded the identity
on every error, so a request that authenticated fine but failed the
action check still reached handleAuthResult with no identity and the
deny path could not audit a requester. Auth now calls
authRequestWithAuthType directly, the same entry AuthPostPolicy uses,
so the resolved identity reaches the error writer; a failed authN
still resolves no identity and stays unattributed. The regression test
now signs a denied request end to end through iam.Auth.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 07:03:01 +08:00
Chris LuandDevin 2864bc0fe8 s3: honor configured session bounds on AssumeRole and LDAP identity (#11478)
* sts: export CalculateSessionDuration

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: honor configured session bounds on AssumeRole and LDAP identity

prepareSTSCredentials hardcoded a one-hour session when the caller
omitted DurationSeconds, so sts.tokenDuration was ignored and
sts.maxSessionLength only clamped explicit requests: asking for 3600s
against a 20m ceiling was rejected while omitting the parameter was
granted a full hour (#11473). The two affected handlers now use the
same default-then-cap calculation as AssumeRoleWithWebIdentity.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: keep MaxSessionDuration through role store copies

copyRoleDefinition rebuilt RoleDefinition field by field and dropped
MaxSessionDuration, so memory-backed role stores silently discarded the
per-role session bound on every write and read (devin on #11478).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* sts: apply per-role MaxSessionDuration to resolved session durations

Review follow-up on #11478 (devin): the role bound only ever applied to
explicit DurationSeconds values — an omitted duration resolved to the
configured default and sailed past a shorter role max on every assume
path.

- capDurationByRole now resolves min(requested||tokenDuration, roleMax),
  so AssumeRoleWithWebIdentity and AssumeRoleWithCredentials cap
  defaults the same way they cap explicit values.
- prepareSTSCredentials caps the calculated duration at the named
  role's MaxSessionDuration, covering the AssumeRole and LDAP handlers;
  self-assumption has no role definition to consult.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: keep MaxSessionDuration through the cached role store

genericCopyRoleDefinition drops MaxSessionDuration the same way
copyRoleDefinition did, so the cached filer role store reads back a zero
maximum and every downstream duration cap is skipped (greptile on
#11478).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* sts: only materialize defaults that pass session duration validation

Review follow-up on #11478 (greptile): materializing an omitted
DurationSeconds into an explicit value could exceed the service's own
input bound (a configured tokenDuration above maxSessionLength) and turn
a previously working request into a validation error.

capDurationByRole now leaves nil anything the service can resolve
better itself, clamps a tightened default at maxSessionLengthSeconds,
and floors a role bound below 900s to the tightest issuable value.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 07:01:51 +08:00
Chris LuandDevin ab95d58b7c s3: keep dedicated object-lock actions pinned during action resolution (#11475)
* s3: keep dedicated object-lock actions pinned during action resolution

A coarse action that already names a dedicated operation (governance
bypass, retention, legal hold, bucket object-lock config) now resolves to
itself before request shape is consulted. Previously a synthetic
DELETE ?versionId authorization request re-resolved to
s3:DeleteObjectVersion, so the bypass check was satisfied by the
delete-version grant alone; with the pin it evaluates
s3:BypassGovernanceRetention as intended.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: cover pinned object-lock actions against competing query params

Locks in the resolution for every dedicated action in the pin set, incl.
the retention and legal-hold shapes carrying versionId.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 21:42:34 +08:00
Chris Luandjsas 2f641a63d6 filer: honor is_moved only from ring member connections (#11456)
* filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks

updateLocalEntry records the RemoteEntry stamp after an upload by writing the
event's entry back with UpdateEntry. The filer deletes every stored chunk
absent from an updated entry, so when the file was rewritten while its upload
was in flight (or the event is a replay), the stale snapshot deletes the
rewrite's chunks: the entry then points at the new fid with no needle behind
it, and the rewrite's own upload fails and is skipped as superseded.

The stamp write now carries WriteCondition IF_CHUNKS_EQUAL over the event's
chunk fids, evaluated by the filer under the path lock. A refused stamp means
the filer moved past this event; the superseding event follows in the log and
stamps the current entry, so the refusal is logged and skipped like a
superseded upload.

Reproduction: weed server -filer plus a weed server -s3 remote, remote.mount,
filer.remote.sync; hold the remote (docker pause) so one upload stays in
flight, rewrite the file through the filer, unpause. Before: the entry's chunk
is 404 on every volume server. After: the stale stamp is refused, the rewrite's
chunk stays live and reads back after a vacuum.

* filer.remote.sync: stamp entries with IF_ENTRY_EQUAL so stale inline content or metadata cannot be restored

The IF_CHUNKS_EQUAL guard compared only the chunk fid multiset, so a
rewrite that touched inline content or metadata alone still compared
equal and the stale snapshot overwrote the live entry. The new clause
compares the whole stored entry against the event's entry under the
same path lock.

* filer: route conditional UpdateEntry to the entry's owner filer

Two filers locking the same path locally could still pass a stale
condition on the non-owner while the owner's entry had moved on. When a
condition or expected_extended precondition is set, forward the request
to the entry's owner the same way conditional CreateEntry does, with
is_moved bounding the hop.

* filer: compare IF_ENTRY_EQUAL against the normalized expected entry

FindEntry grows FileSize to the chunk extent, so a raw event entry with
FileSize still zero failed the condition on an unchanged file and the
stamp was skipped, letting a replay upload the object again.

* filer.remote.sync: classify refused stamps by gRPC status only

A FailedPrecondition substring in an unrelated error would have been
swallowed as a skipped stamp; status.FromError already unwraps.

* remote sync: keep the event entry intact for IF_ENTRY_EQUAL

* filer: honor is_moved only from ring member connections

is_moved is caller-controlled, so a request could set it to skip owner
routing and run a conditional check under a non-owner's lock. Verify the
marker against the peer's connection address and the lock ring members;
an unverified marker is ignored and the request routes like a fresh one.

* filer: refuse unverifiable is_moved at a non-owner, cache ring IPs

Follow-up fixes from review on the is_moved provenance check:

- checkMovedMarker replaces "ignore and re-forward" for markers that did
  not arrive on a ring member's connection. Re-forwarding a claimed hop
  could cycle while rings disagree; instead the request is refused with
  FailedPrecondition unless this filer is the key's owner, in which case
  applying locally is correct anyway.
- ringMemberIPs caches resolved member addresses per ring membership so
  hostname-advertising deployments do not pay a DNS lookup per forwarded
  request; failed lookups are not cached so a DNS blip self-heals.
- DistributedUnlock no longer dereferences the nil response of a failed
  next-hop RPC.

* filer: refuse unverifiable is_moved with PermissionDenied, not FailedPrecondition

A routing refusal is different in kind from a write-condition mismatch:
remote sync treats FailedPrecondition as a stale stamp and skips it, so
reusing that code let a routing failure pass as synced. Owner checks now
also run before the peer-IP lookup so the common accept path does no DNS.

* filer: expire resolved ring member IPs after 5 minutes

A member's hostname can re-resolve to a new IP while its ring address
stays unchanged; caching forever would reject its genuine forwards until
a membership change or restart.

* filer: deduplicate concurrent ring member DNS lookups

At cache expiry, parallel forwarded requests would each resolve every
member hostname serially; singleflight collapses them into one lookup
per ring membership.

* filer: detach the shared ring lookup from the caller's context

The singleflight winner's ctx is cancelled when its request ends; the
shared result would then be an incomplete member list and genuine
forwards denied. The lookup now runs on a detached context with its
own deadline so a canceled caller cannot poison it.

* filer: resolve ring member hostnames in parallel

The shared lookup gave every member one serial budget, so a few slow
resolutions could leave later members out of the cached list and reject
their genuine forwards. Each member now resolves concurrently under its
own detached deadline.

* filer: gather literal member IPs before spawning lookups

A ring mixing IP literals and hostnames raced: the literal appends ran
unlocked alongside the resolver goroutines' locked appends. Split into
two passes so only hostname results share the mutex.

---------

Co-authored-by: jsas <1351492+jsas@users.noreply.github.com>
2026-09-26 19:42:23 +08:00
80fd3635d2 volume: skip TTL last-write scan when it cannot fit its budget (#11472)
* volume: skip TTL last-write scan when it cannot fit its budget

* Update weed/storage/volume_checking.go

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-09-26 19:41:59 +08:00
Chris Lu 7129e1178e s3: evaluate bucket policy before ACL public-read for anonymous requests (#11471)
* s3: evaluate bucket policy before ACL public-read for anonymous requests

AuthWithPublicRead granted anonymous access on a public-read ACL before
consulting the bucket policy, so an explicit Deny (e.g. s3:ListBucket)
was skipped for anonymous callers while still enforced for authenticated
ones. Run the policy engine first: a matching Deny or Allow is honored,
otherwise fall through to the ACL grant as before.

* s3: defer object-level anonymous requests to the handler's policy recheck

Evaluating the bucket policy with a nil entry at middleware time makes
tag conditions like s3:ExistingObjectTag/<key> resolve against missing
values, so a conditional Deny could wrongly block anonymous Get/Head on
a public bucket whose handler recheck would permit it. Object requests
now take the ACL grant and let Get/HeadObjectHandler re-evaluate with
the fetched entry; only bucket-level requests (List, HeadBucket), which
have no such recheck, are decided by the middleware policy verdict.

Reading the bucket config first also refreshes the compiled policy on a
cache miss, so a remotely deleted policy cannot leave a stale verdict
in the engine for nonresident buckets.

* s3: recheck bucket policy before serving directory objects

handleDirectoryObjectRequest runs before the object handlers' policy
recheck, so directory content on a public-read bucket was served to
anonymous callers without any policy evaluation. Evaluate the policy
with the directory entry, matching the recheck the file path performs.
2026-09-26 17:59:57 +08:00
Chris LuandDevin 0f3ba98e11 volume: make volume.scrub report a live needle whose stored id is damaged (#11468)
* storage: scrub live needles' stored id against the index key

scrubVolumeData only compared the needle's stored id for tombstones, so
header damage on a live needle — where the data CRC cannot see it —
passed every scrub mode while reads of that needle kept failing or
serving the wrong key's data. Compare the id for every indexed needle.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: scrub live needles' stored id against the index key (parity)

Mirror the Go scrub fix: compare the stored needle id with the index
key for live needles too, not only for deleted ones.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: cover damaged live needle id in scrub test

The tombstone test proved the index-key check fires for deleted entries;
add the live-needle mirror of Go's TestScrubVolumeDataChecksLiveNeedleId
so a regression in the live path is caught in Rust too.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:10:14 +08:00
Chris LuandDevin 80a26020d7 util: serialize all ViperProxy access so startup cannot hit concurrent map read/write (#11470)
* util: serialize every ViperProxy method; stop promoting unlocked viper calls

ViperProxy embedded *viper.Viper, so only the five declared methods took
the mutex while every promoted call — GetStringMap in backend.LoadConfiguration
was the reported crash — touched viper's maps unsynchronized. `weed server`
starts the volume server (SetDefault writer) and the master (GetStringMap
reader) back to back, and a race build reports the pair on a plain start.

The wrapped viper is now a named field: a method must be declared here to
exist on the proxy, so unsynchronized access fails at compile time rather
than at runtime. Every promoted use in the tree (GetStringMap, GetUint32,
GetFloat64, GetDuration, IsSet, AllKeys, Set) gets a locked wrapper;
NewViperProxy replaces struct literals for local vipers. GetStringMap
deep-copies its result — viper hands back the internal subtree, so
iterating it after the lock is released would race the next writer.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: take the shared lock while LoadConfiguration merges a config file

viper.MergeInConfig rewrites the same maps the proxy serializes; without
the lock a merge can race a concurrent SetDefault or reader exactly like
the reported startup crash.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: deep-copy slice elements in the GetStringMap snapshot

A slice of maps inside the returned subtree still shared the inner maps —
copy elements recursively so nothing the caller mutates is viper's
internal state.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: add the missing AutomaticEnv wrapper used by tests

sse_reader_test reaches it through GetViper(); without the wrapper the
call no longer exists once the viper field stopped being embedded.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: return a fresh slice from GetStringSlice

A stored []string comes back uncast from viper — the backing array is
shared internal state like the GetStringMap subtree, so copy it while
holding the lock.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:09:29 +08:00
5389f61cef volume server: do not finish a GET when the needle CRC mismatches (#11464)
* volume server: do not finish a GET when the needle CRC mismatches

A streamed full-needle read compared the CRC only after every page had been written. Once the response buffer flushed, the client already had a completed 200 and the corrupt bytes. Hold the last page until the checksum matches, and if an earlier page has already been flushed, abort the connection instead of calling http.Error.

Fixes #11459

* volume server: abort partial-content bodies on write error too

The non-Range path drops the unflushed tail and aborts on a mid-body
error; the single-range and multi-range paths still flushed it after
WriteHeader(206) was committed, delivering corrupt bytes as a complete
body.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: assert the started 200 is aborted in the write-error test

The test previously returned on any request error, so it passed without
verifying the abort. It now asserts the client got the committed 200
headers and then a failed body read. Also trims comments.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:00:22 +08:00
Chris LuandDevin 8ad2f29e3e shell: let volume.deleteEmpty drop volumes with no live needles (#11437)
* shell: let volume.deleteEmpty drop volumes with no live needles

The candidate check only accepted a .dat at superblock size, so a volume
whose every needle was deleted still had to be vacuumed first — minutes
of compaction to rewrite bytes that were all garbage anyway. FileCount
counts every indexed entry and DeleteCount every entry made garbage by
overwrite or delete, so FileCount <= DeleteCount means nothing live
remains and the volume can be unlinked directly. The quietFor guard is
unchanged.

* volume server: add only_garbage VolumeDelete guard

VolumeDelete(only_empty) refuses every volume that ever held data, so a
volume whose needles are all deleted could only be removed after a
vacuum rewrote it. The new only_garbage flag deletes only when the byte
counters show nothing live: DeletedSize covering all of ContentSize, the
same all-garbage state vacuum measures. Byte counters are used because
the file/delete counts drift on index reload.

* rust volume: mirror only_garbage VolumeDelete guard

Same check as the Go server: a volume deletes under only_garbage when
its deleted bytes cover all content bytes. The grpc handler rejects
before the store drops the volume from its map, since destroy errors
after removal would still unmount it.

* volume delete: let either enabled check pass, keep onlyEmpty on the wire

An upgraded shell sending only_garbage to a pre-upgrade server would be
read as an unconditional delete (field ignored, only_empty false). The
request now keeps only_empty set so old servers check emptiness and
refuse, while new servers delete when either check passes.

* volume.deleteEmpty: skip remote-backed and protected read-only volumes

A remote-tiered replica shares its cloud object with the other replicas,
so keepRemoteData=false on one delete removes data they still reference.
Protected read-only volumes are quarantined or under maintenance, which
is exactly when a replica should not be dropped.

* volume delete: validate guarded copies across disks before deleting

* volume delete: hold copy locks across guarded validate-and-delete

CheckVolumeDeletable released each copy's locks before Destroy ran, so a
write landing on a later copy between the two passes refused its destroy
after earlier copies were already removed. Pin every copy's
dataFileAccessLock (and its location's volumesLock) across validation and
removal so a refused delete leaves all copies intact.

* volume delete: send deleted-volume notices after releasing locks

A blocking send on a full DeletedVolumesChan under volumesLock can stall
the heartbeat loop that drains it while it waits on the same locks.
Collect the notices under the lock span and send after release.

* pb: restore generated-file cosmetics to match the repo's protoc version

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 15:53:39 +08:00
Ilia Demianenko c58bd0dfd3 s3: honor assignment fsync in UploadWithRetry (#11449)
* fix: honor assignment fsync in UploadWithRetry

* Tests feedback
2026-09-26 12:01:52 +08:00
975cec9228 s3api: exclude marker part in listObjectParts pagination (#11463)
* s3api: exclude marker part in listObjectParts pagination

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* s3api: guard listObjectParts marker boundary and enhance pagination test

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* s3api: fold in review feedback from the parallel #11462 fix

Same core fix; this adds the explanatory comment, tightens the overflow
guard to math.MaxInt64, makes the fake filer sort entries like a real
listing, and adds the marker-exclusivity assertions alongside the
pagination walk.

Co-authored-by: yi111 <yi111@users.noreply.github.com>

---------

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: yi111 <yi111@users.noreply.github.com>
2026-09-26 11:58:23 +08:00
Chris LuandDevin f31a026b2a master,filer: fix lock ring poisoning after leader change (#11453)
* cluster: never broadcast an empty lock ring

An empty member list is never a usable ring state, but a delayed
RemoveServer on a former leader can fire after the new leader already
broadcast the recovered ring. That late broadcast carries a newer
wall-clock version, so clients accept the empty ring and permanently
reject the good one.

Skip the broadcast entirely when the member list is empty, keeping the
last non-empty snapshot for reconnecting clients.

* cluster: periodically rebroadcast the lock ring

Ring updates are purely event-driven, so one lost or poisoned update is
permanent until the next membership change — with a single filer that may
never come. Re-arm a per-group timer after every broadcast so the current
leader keeps re-sending the ring; clients reject nothing newer than their
last accepted version, so a re-sent snapshot always heals a stale view.

* filer,s3api: reset the lock ring on master change

Ring versions are per-master monotonic — each master stamps wall-clock
nanoseconds — so a late high-version update accepted from a former leader
makes the new leader's snapshot look stale forever. Detect a leader
change across the reconnect gap (currentMaster is cleared between
attempts, so remember the last served master) and reset the ring to
bootstrap state so the new leader's view always applies.

* cluster: fail lock acquisition when no lock server exists

retryUntilLocked loops forever, so a filer reporting an empty lock ring
wedges every append write indefinitely. Bound only the "no lock server
found" case — ordinary contention is still waited out since the holder
releases eventually. The constructors now return nil on failure: the
filer append path and S3 object writes fail fast, while mounts degrade
to their existing lockless mode.

* cluster: reset only the ring version on master change

Ring versions are per-master monotonic, so a version gate reset is all a
leader change needs. Clearing the whole ring made every filer its own
write owner until the next update and dropped the prior-owner window for
keys the new leader remaps; the last ring now keeps routing until the
new leader's snapshot transitions off it.

* cluster: skip redundant ring installs and defer rebroadcasts

An unchanged member list now only bumps the accepted version instead of
installing a snapshot: periodic rebroadcasts no longer fire the
topology-change callback or restart the prior-owner window. And a
rebroadcast that lands inside a membership stabilization window yields
to the pending timer rather than publishing an intermediate ring.

* cluster,mount: bound lock unavailability, fail ops that cannot lock

Only 'lock already owned' contention retries without bound now; every
other failure — no lock server, or a dead ring member refusing
connections — shares the same unavailability budget, so a ring naming
departed filers can no longer hang a lock forever. Mount open-write,
create, and rename fail with EAGAIN when the required lock cannot be
acquired instead of proceeding without cross-mount serialization.

* cluster: check pending stabilization inside the broadcast critical section

rebroadcast released the mutex between the pending-timer check and
nextBroadcastUpdate, so a membership change arriving in the gap could arm
a stabilization timer while the rebroadcast emitted an intermediate ring.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: acquire path locks before mutating create/rename state

Create took the DLM lock only after the filer create, so a lock failure
returned EAGAIN with an eagerly persisted file left behind. Rename marked
source handles renamed before acquiring locks, so a failed acquisition
left them suppressing old-path flushes for a rename that never happened.
Both now take the locks first; the create's lock is released again if the
entry race loses to another creator and AcquireHandle takes over.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: keep the old-path lock when rename lock migration fails

The migration stopped the handle's lock before acquiring the replacement,
so a nil result left the handle writing with no lock at all. Acquiring the
new-path lock first means failure keeps the existing lock instead of
reporting success with serialization dropped.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: skip new-path rename lock when a handle already holds it

A target file open for write on this mount already carries a lock on
newPath; the lock manager does not grant a second lock to the same
owner, so the rename would wait on itself until the handle closed.
Also avoid locking twice when old and new paths coincide.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: hand the rename's target lock to the migrating handle

The rename holds a lock on newPath for its duration, so the response
migration's fresh acquisition waited on that same lock until the handle
released — under fhLockTable, blocking the handle's own close. Adopt the
rename's lock directly; nested move responses still acquire their own.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: move the replaced target's lock to the renamed handle

When the target path was already locked by an open handle on this
mount, the migrated source handle kept only its stale old-path lock —
the target's close would then release the last lock on the new path
while the renamed handle was still open. Adopt the replaced handle's
lock instead.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: stop the handle lock inside the fh lock on release

ReleaseHandle stopped fh.dlmLock before taking the fhLockTable slot, so
a rename migration holding that slot could still observe and adopt a
lock that was already stopping. Stopping under the fh lock makes the
transfer serialize against the release.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: claim the replaced target's lock for the renamed handle

When the target path is already locked by an open handle on this mount,
adopting it at migration time keeps the renamed path protected after
that handle closes, without waiting on a lock this mount already holds.
If the handle was released mid-migration the claimed lock is stopped,
and a fresh acquire covers the case where it was already gone.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: claim the target handle's lock before the rename runs

Skipping the new-path lock when a handle already holds it let that
handle's close release the lock mid-rename, leaving the path unguarded
until the response migrated it. Take over the lock at check time and
hold it for the rename's duration: the response adopts it for the
migrating handle, or it returns to the target handle / is released on
failure. The target handle lookup also falls back to the entry's stored
inode for a forgotten path mapping.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: read handle locks only under the fh lock during rename

The loose dlmLock reads raced ReleaseHandle, which now mutates the lock
inside the handle lock; check and claim it under the same hold.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 11:57:00 +08:00