Commit Graph
9747 Commits
Author SHA1 Message Date
Chris Lu da087f77b3 mount: stop a replaced rename destination from flushing over the rename (#10965)
* mount: stop a replaced rename destination from flushing over the rename

Rename replaces whatever the destination held, which deletes that entry, but
only the source handle was told. A handle still open on the replaced entry
went on flushing its metadata under that name, and on Windows -- where the
close carrying the flush runs after the application's CloseHandle has already
returned -- the flush landed after the rename and put the destination's old
content back:

    dir Rename old_entry:{name:"src"} new_entry:{name:"dst" ... inode:...3416}
    doFlush /dst fh 1521468582993181449
    /dst saveToStorage 1,6872462993 [0,3)
    flushMetadataToFiler /dst inode 11939747521756968515
    InsertEntry /dst

The next read of the destination returned the content the rename was supposed
to replace. Unlink already handles this with markHandleDeleted, which raises
the flag under the handle's flush lock so a flush already writing finishes
first and any later one sees it; a rename that replaces an entry deletes it
just the same, so it now does likewise.

Verified on the Windows runner: TestRenameOverExisting 300/300, where the same
loop reproduced the corruption twice without this.

* test/winfsp: say which layer kept a renamed-away name

The failure only reported the stat. Which layer answered narrows the search a
lot: a listing reads no per-path cache, the mount's own forgets within a
second, and a name that survives both is still in the meta cache.

* mount: keep the destination barrier honest when the rename does not happen

Two gaps in the barrier the previous commit put in front of a replaced rename
destination:

The flag was raised before the filer rename, which can still fail. The
destination then stays exactly where it was, with its handle marked deleted
and its dirty metadata silently dropped from then on, so a rename that
returned an error has to put the flag back.

The handle was only found through the path mapping, which Forget drops while
the handle is still open. The source side already falls back to the inode the
entry carries; the destination now does the same, off the entry the sticky-bit
check had already loaded.

* mount: let only the caller that raised a delete mark lift it

Restoring the destination handle after a failed rename cleared isDeleted
outright, so an unlink that marked the same handle in between lost its mark and
a later flush could write the unlinked entry back.

Every raise of the flag already happens under the handle's flush lock, so
counting them there is enough to tell one caller's mark from another's: the
rename lifts only the mark it made itself.

* mount: drain the destination flush before marking it deleted

A flush already queued for the destination belongs to the entry as it stands.
Marking first meant the drain waited on a flush that then skipped its metadata
as deleted and released its handle, so a rename that failed afterwards had
nothing left to restore and the queued update was gone, its chunks orphaned.

Draining first lets that flush finish as itself, before the rename has taken
anything away.
2026-08-26 08:51:37 -07:00
Chris Lu eb3bbfeb1f filer: apply the path's storage rule TTL on every write path (#10963)
* filer: cover the storage rule TTL on the object transaction write path

An object written through ObjectTransaction used to land with ttlSec 0
even under an fs.configure TTL rule, while the same object written
through CreateEntry got the rule's TTL. Guard the shared stamping so the
two paths cannot drift apart again.

* filer: apply the path's storage rule to an appended entry

AppendToEntry resolved the storage option from the path - so its chunks
land on a TTL volume under an fs.configure TTL rule - but never stamped
the rule's TTL on the entry it creates, leaving an entry that outlives
its data. Route it through applyStorageDefaultsToEntry, which now feeds
the entry's own TTL into the option so the placement an existing entry's
appended chunks get is unchanged.

* filer: apply the path's storage rule to a completed TUS upload

The PATCH path resolves the storage option from the target, so a TUS
upload into an fs.configure TTL prefix writes its chunks to a TTL volume,
but completion built the final entry with ttlSec 0 - the entry outlived
the data it pointed at. Stamp it through applyStorageDefaultsToEntry,
which also subsumes the hand-rolled read-only check and supplies the
rule's name-length limit.

* filer: apply the destination's storage option TTL to a copied entry

The copy handler re-uploads the source's chunks under the destination's
storage option, so a copy into an fs.configure TTL prefix already lands
its data on a TTL volume. The entry, though, carried the source's ttlSec
- 0 for a source outside the prefix, or the source's own TTL where the
two rules differ - so it never expired with the data it pointed at. Take
the TTL from the same option the chunks were placed with, after the
data-only copy has restored the destination's metadata.
2026-08-26 08:49:25 -07:00
Chris Lu a02c0024e5 master: cap the reported capacity at what the disks hold (#10960)
* master: cap the reported capacity at what the disks hold

Statistics reported max volume count times the volume size limit, which is
how many volumes the cluster is allowed to place, not how much space it has.
A cluster given far more slots than its disks can fill reported a capacity it
could never reach -- 65536 slots at 30GB read as 1.9PB on a 460GB disk -- and
the number never moved, since writing data changes neither the slot count nor
the size limit.

The volume servers already report each filesystem's total and free bytes in
their heartbeats, so bound the answer by what they say is left.

* mount: keep the last known sizes when filer statistics fails

A failed Statistics call returned before df's answer was filled in, so a
mount whose filer or master was briefly unreachable reported an empty
filesystem rather than the sizes it already had.

* master: drop the disk ceiling when a volume server does not report

A cluster part way through an upgrade has volume servers that predate the disk
bytes in the heartbeat. Summing only the ones that answered left the quiet
server's free space out of the total, and the server holding the room is
exactly the one that could make the cluster read as full.

Answer with the disks only when every one of them reported.
2026-08-26 00:12:56 -07:00
Chris Lu 7658305c76 mount: name the disk after the mounted path (#10958)
* mount: name the disk after the mounted path

Finder and Explorer labelled every mount with the filer address, so two
mounts from one filer were indistinguishable. Use the mounted path's last
segment, the way df already shows it, and keep the filer address only for
a whole-tree mount.

* mount: let a given mount option override the default

The options from -o were placed before the ones this mount derives, so
a volname or iosize given on the command line lost to the derived value.
Append them last, matching the Windows adapter.

* mount: document what labels the disk
2026-08-25 22:56:33 -07:00
Chris Lu 627b5e9d59 shell: parse every collection filter the same way (#10955)
* worker: move the collection filter parser into weed/util/wildcard

The parser sits beside the volume-list filtering it was written for, in
weed/plugin/worker, which imports weed/shell — so the shell commands that
parse the same filter three other ways can never call it. Move it down to
weed/util/wildcard, next to the comma-separated wildcard helper it already
replaced, leaving the behavior unchanged.

* shell: parse every collection filter the same way

The shell parsed a collection filter three ways: compileCollectionPattern
compiled one regex for ec.encode, ec.decode, volume.balance and the tier
commands; volume.list and volume.deleteEmpty matched a single wildcard; and
volume.tier.move, volume.fix.replication and volume.configure.replication
called filepath.Match on their own. None of them took a list, so
"ec.encode -collection=a,b" selected nothing, the same way the admin UI did.

They all go through the shared matcher now: a comma-separated list of names,
"*" and "?" wildcards, "_default" for the collection with no name, and regex
entries. The one thing that stays per-command is what an empty value means -
every collection for -collectionPattern, the unnamed collection for the ec
and tier -collection flag - so compileCollectionPattern keeps that mapping.

The matchers are compiled once per command instead of once per volume, and a
regex entry now has to match the whole name unless it anchors itself, so
-collection=bucket no longer picks up mybucket2.

* shell: keep dots in collection names, and commas inside a regex

A dot no longer marks an entry as a regex, so a collection named "my.bucket"
matches itself and not "my-bucket" - the difference decides which volumes
volume.deleteEmpty and volume.tier.move touch. A dot still counts when it is
quantified, so "bucket.*" stays a prefix regex.

The comma split also leaves alone the commas inside a character class or a
repetition count, so "bucket[0-9]{1,3}" stays one entry instead of becoming
two broken fragments.

* shell: let a regex entry match its own spelling

A collection named after regex syntax, say "logs(2024)", was unreachable:
the entry compiled to a pattern that matches "logs2024" instead. Match the
entry verbatim as well, so naming a collection always selects it, whatever
characters it holds.

* shell: reject a collection filter that names no collection

A value of "," parsed to no entries and then matched every collection, so a
typo widened ec.encode or volume.deleteEmpty to the whole cluster. Only a
genuinely empty filter means "all collections"; anything else has to name one.

* shell: keep commas inside a regex group out of the entry split

The split already left alone the commas inside a character class or a
repetition count, but not the ones inside a group, so "bucket(foo,bar)"
was cut into two fragments that no longer compile.

* shell: cover escaping a collection name that is not a regex

A name like "logs(2024" does not parse as a regex on its own; escaping it,
"logs\(2024", reaches it. Pin that so the escape hatch does not regress.

* shell: split entries only on commas inside a closed regex construct

An unmatched "{" or "[" made the splitter swallow every comma after it, so
"foo{bar,videos" became one entry that matches neither collection - the
silent no-op this filter work exists to remove. A construct now has to close
before its commas stop separating entries.

* shell: skip character classes while scanning a regex group

A ")" inside a class is a literal, so "(a[)],b)" ended its group early and
split into two fragments that no longer compile.

* shell: cover escaping a comma inside a collection name

A comma separates entries, so a name holding one is reached by escaping it.

* shell: follow the regexp parser when scanning a character class

A "]" leading a class is a member of it, and a POSIX class such as
"[:alpha:]" carries its own "]", so stopping at the first one cut a valid
filter like "(a[]),],b)" into fragments and rejected it.
2026-08-25 18:03:52 -07:00
Chris Lu 368b2035b2 s3: deny anonymous access when the identity config loads no identities (#10954)
* s3: deny anonymous requests when the identity config loads no identities

Naming a config file is the operator asking for authentication. A file that
yields no identity - an unpopulated secret mount, or a mistyped top-level key
the proto parser silently drops - left the gateway open to every anonymous
caller: ListBuckets returned 200, and anonymous PUT could create buckets and
write objects.

* s3: name the unknown top-level keys in an identity config

The proto parser discards what it does not recognise, so a mistyped
"identites" loads as an empty config. Naming the dropped keys at startup turns
the resulting lockout into a one-line diagnosis.

* s3: isolate the auth-enforcement tests from AWS environment credentials

* s3: use a singular "identity" as the unrecognised-key example

Codespell rejects the misspelling the example used.

* s3: cover the empty identity config alongside the unrecognised key

* s3: cover a config file whose body is an empty object
2026-08-25 15:31:32 -07:00
Chris Lu e482e67971 admin: accept a list of collections in the task collection filter (#10953)
The collection filter was parsed twice with two syntaxes: the master-side
volume listing compiled the whole string as one regex, while EC encode and
EC balance detection split it on commas and matched each entry as a
wildcard. A volume had to pass both, so "collection-a,collection-b" matched
nothing (no collection is named that), and the ALL_COLLECTIONS sentinel,
which the master side skips, dropped every volume at the task side.

Parse it once, in one place: a comma-separated list where an entry is a
name with optional * and ? wildcards, or a regex when it carries regex
syntax. A regex entry now has to match the whole name unless it anchors
itself, so listing a collection no longer picks up its longer namesakes.
2026-08-25 15:13:06 -07:00
ef4c9d9178 filter volume by local or remote storage name (#10946)
* filter volume by local or remote storage name

Signed-off-by: lou <alex1988@outlook.com>

* fix SelectsEverything

Signed-off-by: lou <alex1988@outlook.com>

* keep the proto sync out of this change

The branch copied weed/pb/*.proto over their seaweed-volume and Java
counterparts and regenerated every .pb.go with a different protoc and
protoc-gen-go-grpc. DiskStatus.error arriving that way broke the Rust
build, and the rest is toolchain churn in files this change has nothing
to say about.

---------

Signed-off-by: lou <alex1988@outlook.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-25 13:05:33 -07:00
Chris Lu 70c3adb983 volume: stop read-only volumes from pinning .idx and .sdx (#10950)
A read-only or cloud-tiered volume loads a SortedFileNeedleMap, which held
both its .idx and its .sdx open for the life of the process. On a server with
~600K tiered volumes that is 1.2M descriptors before a single read, enough to
exhaust the fd limit and take the listeners down. The .dat is not the problem:
a tiered volume serves it from the remote backend.

Neither index file is needed except while a lookup is in flight, so borrow them
from a bounded process-wide pool instead. An idle volume now holds zero
descriptors; a busy one keeps its handles hot rather than paying an open() per
needle. Reads borrow O_RDONLY, so a volume on a read-only mount answers lookups
that previously failed at load. Sync tracks whether a tombstone was appended,
which also drops the fsync-per-volume storm at shutdown.
2026-08-25 10:23:00 -07:00
Chris Lu b77431c142 master: stop hintless small-file assigns from marking volumes full (#10944)
* master: estimate a hintless assign's size from the volume's average file size

An assign that carries no dataSize hint charged a flat 1MB per file id
against the volume's effective size. A small-file workload overpays by
orders of magnitude: bulk-writing 4KB files marks volumes holding a few
hundred MB of real data as crowded and then full, so the master grows
unnecessary volumes and, once every volume is spuriously full, fails all
assigns. Estimate from the volume's own average file size instead, and
keep the 1MB fallback only for volumes with no history.

* master: decay pending assign sizes for volumes gone quiet

The decay that corrects pending assign estimates runs only when a
heartbeat reports the volume, and a heartbeat only reports a volume
whose content changed. A volume held out of the writable list takes no
writes, so once inflated estimates mark every volume full, nothing is
ever reported again, nothing decays, and the cluster refuses all writes
until a restart. Run the decay from the master's periodic loop for
volumes no heartbeat has reported within two pulses, feeding the last
reported size back through the same path an unchanged heartbeat would
take.

* master: trim the comments on the assign size estimate

* master: keep the periodic decay out of the replica-dedup window

UpdateVolumeSize ignores a report arriving within two seconds of the last
one, so replicas of the same volume do not each halve the pending
estimate. The periodic decay went through the same path and stamped that
window, so a real heartbeat landing right behind it was dropped along
with its reported size and compact revision. Only a volume whose content
changed is reported at all, so nothing would send that size again and
the master kept a stale one. Let the dedup window belong to volume
server reports alone.

* master: let the decay read the size record under the lock it mutates

The periodic decay picked its volumes under a read lock and replayed
them under a write one, carrying the size it had read across the gap. A
heartbeat landing in between was rolled back: the replay wrote the older
size and compact revision over the fresh ones, and a compaction report
lost that way is never resent, since only a volume whose content changed
is reported. The decay has no size of its own to contribute, so it now
reads the record under the same lock it mutates.

* master: let a heartbeat that beat the decay stand for the cycle

The decay chooses its volumes under a read lock and applies them under a
write one. A heartbeat landing in that gap already did the halving the
cycle owed, so applying the decay on top of it halved twice and forgot
pending bytes the volume has not written yet - the double-halving the
replica-dedup window exists to prevent. Both callers now give way to a
report already handled for this cycle; only a real report still advances
lastUpdateTime, so a quiet volume keeps decaying every pulse.

* master: keep genuinely full volumes out of the decay pass

A volume the disk really did fill keeps its fullSince set for good, so it
was selected every pulse for a decay that cannot help it: UpdateVolumeSize
refuses to recover a volume whose reported size is at the limit, and
replaying a size that cannot move leaves the record as it found it. Full
and quiet is the ordinary resting state of a cluster, so this was most of
the pass, taking the layout write lock away from the heartbeats to do
nothing. On a million tracked volumes with a hundredth of them phantom-full
it costs ten thousand write locks a pulse instead of a million.

* master: put the stale-replay test back on the path it guards

Giving the decay the dedup window left this test short-circuiting there,
so it no longer reached the locked read it was written for and passed
with that read removed. Age the record past the window, which is the only
case where reading it under the lock is what saves the report.
2026-08-25 10:16:48 -07:00
Chris Lu 50b388771a s3: stop one abandoned request from cancelling every concurrent upload (#10948)
* grpc: a non-cancellable context is no evidence of a stale channel

shouldInvalidateConnection only invalidates on Canceled/DeadlineExceeded
while the context handed to WithGrpcClient is still live, so that an RPC
timing out on its own does not close the shared cached ClientConn and
cancel every other in-flight RPC on it. context.Background()/TODO never
expire, so Err() stays nil forever and that guard always answered
"invalidate" - and Background is what almost every caller passes, the S3
gateway included.

One S3 request whose RPC rode an abandoned HTTP request context therefore
closed the shared filer connection, and every multipart part in flight
died with "the client connection is closing", surfacing to the client as
400 InvalidRequest.

Only a cancellable context bounds an RPC attempt, so require one before
reading it. A genuinely stale channel (a peer restart behind a stable L4
endpoint) surfaces as Unavailable, which invalidates on its own branch.

* grpc: a bystander of a connection teardown is not a stale-channel witness

gRPC raises ErrClientConnClosing locally, before an RPC reaches the wire,
when this process has already closed the ClientConn. Every caller that
touches a channel during another goroutine's teardown gets it, so reading
it as a stale-channel signal lets one teardown re-arm itself across the
whole herd of callers it just cancelled.

The cached-connection version check keeps those callers from closing a
replacement channel, but the streaming path invalidates by address alone
and has no such guard.

* grpc: end a stream without dropping the peer connection under it

A streaming caller gets its own ClientConn, but on any error it also drops
the cached non-streaming ClientConn every request handler shares with that
peer, to recover a peer restart hidden behind a stable L4 endpoint. Any
error includes the ordinary ones: a metadata subscription that reached its
stop point, a follow callback that refused an event, a caller that gave up.

The S3 gateway follows filer metadata on such a stream and reconnects
forever, so each ordinary end of it cancelled every S3 request in flight
against the filer. Drop the shared channel only for errors that say the
peer went away, which is what invalidation is for.

* test: close the connections the cascade tests leave cached

Each test swaps in a fresh connection cache and restores the previous one,
dropping its own entries without closing them, so the ClientConn's
transport and reconnect goroutines outlive the fake filer they dialed.

* grpc: say why ErrClientConnClosing's deprecation notice does not apply

It points at codes.Canceled, which is the code this function exists to
disambiguate. Only the message distinguishes a teardown a caller merely
walked into, so the sentinel stays.
2026-08-25 10:15:47 -07:00
Chris Lu 44115c1051 filer: stop TUS uploads from turning into garbage (#10945)
* filer: store TUS sub-chunks through the regular chunk writer

A TUS sub-chunk was written with one assigned file id, retried up to
three times against that same id, and abandoned on failure: an attempt
that had landed on some replicas left a needle no session record and no
entry ever references, unreclaimable by vacuum.

dataToChunkWithSSE, which the regular write path uses per chunk, assigns
a fresh file id per attempt and hands back the file ids of failed
attempts, which are now freed the way the regular write path frees them.

* filer: retry a chunk write on a fresh volume when the server 5xxs

The filer's chunk writer assigns a fresh file id per attempt but only
retried transient network errors, so a volume filling up and turning
read-only mid-write failed the whole request even though the very next
assignment would have landed elsewhere. Every other write client already
routes this through ShouldReassignUpload; the filer's own write path now
does the same, for regular uploads and TUS sub-chunks alike.

* filer: export the chunk deletion queue

The filer test harness in weed/server builds filer.Filer as a struct
literal, so any code path reaching DeleteChunks dereferenced a nil
queue. Exported like the neighboring DeletionRetryQueue so the harness
can arm it.

* filer: complete a TUS upload whose chunk records overlap

A PATCH retried while its predecessor was still storing a sub-chunk -
a proxy timeout with an immediate retry is enough - records the same
range twice. HEAD computes Upload-Offset as the covered watermark and
reported the upload fully received, but completion demanded exactly
adjacent records and failed every attempt: the client concluded success
from offset == length, no entry was created, and the session eventually
expired, turning the entire upload into deleted needles for the vacuum
to chew through.

Completion now validates gapless coverage with the same watermark HEAD
uses. A record extending coverage joins the entry - the read path
resolves partial overlaps by ModifiedTsNs, and the raced copies carry
identical bytes - while a fully covered duplicate is freed once the
entry lands.

* filer: allow one mutating TUS request per session at a time

Nothing stopped two PATCHes from writing the same range concurrently:
both loaded the same offset, both passed the conflict check, and both
recorded their sub-chunks. A client whose request timed out in a proxy
retries immediately while the server side is still storing the buffered
sub-chunk, which is exactly that race.

A session now accepts one PATCH or DELETE at a time, the way tusd locks
uploads; a concurrent one is refused with 423 Locked, which TUS clients
retry, and HEAD keeps answering so progress polling is unaffected. The
chunk state is loaded under the claim, so a retried PATCH sees every
record its predecessor left and conflicts cleanly instead of duplicating
data.

* test: cover a TUS PATCH raced by its own retry

Stalls a PATCH mid-body over a raw connection, retries the same range
while it is in flight, and expects the retry refused with 423 Locked;
the upload then resumes from the reported offset and the final content
must be intact.

* filer: never free a TUS duplicate the entry still references

Coverage is computed from ranges, so a record fully covered by another
is treated as a duplicate no matter which needle it names. A malformed
record naming a file id the entry keeps would have had that needle freed
right after the entry landed - the corruption this change set exists to
stop. The duplicates are now freed in one batch, skipping any file id
the entry references; their records go with the session directory.

* test: bound the raw TUS connection reads

http.ReadResponse on the stalled PATCH's connection blocked until the
whole go test timeout if the filer never answered.

* filer: free the needles of chunk write attempts a retry replaced

A volume server stores the needle locally and only then fans out to the
replicas, so a replication failure 5xxs with the data already written.
Each attempt assigns its own file id, so once a later attempt lands
elsewhere nothing references the earlier ones: the caller only sees the
chunk that succeeded, and the failed ids were dropped.

They are now freed the way the caller frees them when the whole write
fails. Retrying on a 5xx makes this reachable on every read-only or full
volume, which is exactly the condition that filled the reporter's
volumes.
2026-08-25 09:24:51 -07:00
Chris Lu 68f0793b6f mount: register UNC mount points as WinFsp network file systems (#10943)
A \\server\share -dir was passed to WinFsp as a plain mount point, which
treats it as a directory path on an actual remote server and fails. Turn it
into the VolumePrefix option instead, so the mount registers with the WinFsp
network provider: the UNC path is then reachable from every logon session,
which a drive letter mounted from a service is not, and each user can map
their own drive letter to it.
2026-08-25 01:28:50 -07:00
Chris Lu b3be2f5449 filer.backup, filer.sync: stop sharing resume checkpoints across destinations (#10934)
* filer.backup: key the checkpoint by source path and sink destination

The checkpoint id hashed only sink name + directory, so two backups to
different buckets or endpoints sharing a directory layout advanced one
checkpoint: whichever job was running pushed the shared offset forward,
and a stopped or failing job later resumed from the other's position,
silently skipping changes. Backups of different source paths to the same
destination shared a checkpoint the same way.

Each sink now reports a destination identity (endpoint or account,
bucket or container, directory) and the checkpoint is keyed by the
source path plus that identity. Reads fall back to the historical
name+directory key when the new key has no value, so existing backups
resume where they left off; writes go only to the new key.

* filer.sync: include the target path in the offset key

The offset stored on the target filer was keyed by source path and
source filer signature only, so two syncs from the same source cluster
and path to different directories on the same target cluster advanced
one shared checkpoint, and the slower one could resume past events it
never applied. The target path now participates in the key; "/" keeps
the historical form, and a sync with a non-root target path falls back
to the historical key once when its own key has no value yet.

* join checkpoint key fields with NUL so they cannot alias

A path or configuration value spelling out the separator could
concatenate two different field tuples to the same checkpoint key.
NUL cannot appear in a CLI path argument or any sane configuration
value, making the encoding injective.
2026-08-24 19:30:20 -07:00
Chris Lu 4a2879abad admin: show a copyable S3 object URL in the bucket file browser (#10933)
* admin: offer copyable S3 object URLs in the bucket file browser

* admin: hide object urls when the bucket type lookup fails

* admin: ignore an s3.public_endpoint that is not an absolute http url

* mini: build the seeded s3 endpoint with JoinHostPort for ipv6

* admin: reject a query or fragment in s3.public_endpoint

* mini: drop the seeded s3 endpoint when a later run disables s3

* admin: reject userinfo and bare delimiters in s3.public_endpoint, redact the warning

* mini: pass its s3 endpoint as an admin option instead of mutating viper

* admin: keep the rejected s3.public_endpoint value out of the log
2026-08-24 19:29:01 -07:00
Chris Lu 2a70532d0d s3: log each request at -v=2 (#10931)
* s3: log each request at -v=2

* s3: quote requester and path in the access log line

* s3: record the post-policy signing identity as the requester
2026-08-24 18:39:30 -07:00
Chris Lu d2c470af1b S3: commit SSE GET status only after the first read succeeds (#10935)
The SSE streaming path kept writing 200/206 from filer metadata before
fetching or decrypting anything, so a missing needle or failed decrypt
setup surfaced as a broken 200 body. Same deferral as the plain path:
the status commits on the first body write, and every failure before
that returns to the handler for a clean S3 error response.
2026-08-24 18:36:46 -07:00
Chris Lu d9d5fab35b S3: commit GET status only after the first read succeeds (#10930)
streamFromVolumeServers wrote the 200/206 status from filer metadata
before any byte had been fetched from a volume server, so a missing or
corrupted needle surfaced as a broken 200 body and the request metrics
recorded a success. Defer the status commit to the first body write: a
failed first read now returns a clean 500 before headers, while the
wire timing of successful responses is unchanged since net/http buffers
the status line until body bytes arrive anyway.
2026-08-24 16:14:52 -07:00
MaratKarimovandMarat Karimov a3afe4460b tarantool: fix upsert data corruption and missing context propagation (#10926)
Co-authored-by: Marat Karimov <karimov_m@inbox.ru>
2026-08-24 15:17:08 -07:00
Chris Lu 863fec6c3f S3: let a key that is a prefix of other keys be an object (#10912)
* filer: keep the sentinel when CreateEntry reports an update failure

CreateEntry flattened the error UpdateEntry wraps, so errors.Is stopped
matching and ErrExistingIsDirectory and ErrExistingIsFile never reached
the S3 mapper, which answered a retryable 500 instead.

* s3: let a key that is a prefix of other keys be an object

S3 keys are flat, so "a/b" and "a/b/c" are independent objects that
coexist in either write order. The filer stores a key as a path, so one
of them has to live on the directory the other is nested under.

Writing the nested key first refused the prefix key outright. Writing it
second promoted the file to a directory, which kept its data but lost the
key: an empty object left nothing to recognise it by and disappeared, and
one with data listed under a trailing slash it never had.

Mark the directory that carries such a key, and write the object onto it
when the path is already a directory. The mark makes an empty prefix
object visible to listings and readable by GET and HEAD, keeps the empty
folder cleaner off it, and lists it under the key it was written with.
Deleting the key strips the mark back off along with the data.

* filer: keep a TTL off a directory that stands for an object

An expired entry is deleted a row at a time, so expiring a directory
removes it and leaves everything under it unreachable. Promoting a file
to a directory carried its TTL across, and a promoted file is exactly the
one that has keys nested under it.

Drop the TTL on promotion, and leave one an older build wrote alone. The
lifecycle worker still expires the object, through the delete that leaves
the directory behind.

* s3: delete the null version of a key other keys are nested under

The routed delete cannot remove an entry that other keys live under, and
answered a retryable 500 rather than falling back to the lock path the
unversioned delete already falls back to. That path then looked the entry
up under the bucket with the whole key as its name, so the demote wrote it
back one directory too high and failed as not found.

Fall back on any non-precondition error, and split the key before deleting
it. Trailing-slash directory markers with children reach the same delete.

* filer: keep the sentinel when MkFile and Mkdir report a create failure

Same flattening one layer out: every mkFile caller lost the sentinel, so
a CopyObject onto a key that other keys are nested under answered a
retryable 500 where a PutObject of the same key answers 409.

* s3: copy and rename a key that other keys are nested under

Such a key is stored on the directory those keys live in, and copy and
rename both refused it: the source lookup maps every directory entry to
NoSuchKey, so a key a plain GET serves could not be copied or moved, and
the destination side refused it as a directory conflict.

The source is read through a view of the entry as the object it names.
The destination is written the way a PutObject of that key writes it. A
rename at either end copies the object's own data across and strips it off
the source key rather than going through AtomicRenameEntry, which moves a
directory by moving everything under it - the nested keys are not part of
what is being renamed.
2026-08-24 15:10:34 -07:00
Chris Lu 46ce2c45a2 mini: reserve the admin gRPC port instead of binding it late (#10928)
* mini: reserve the admin gRPC port instead of binding it late

Port selection probes every port with a throwaway listener and closes it.
Master, filer, volume and S3 bind a moment later, but the admin waits for
all of them first and only then binds its worker gRPC port, roughly two
seconds in. That port defaults to the admin http port + 10000, which lands
inside the Linux ephemeral range, so one of the cluster's own outgoing gRPC
dials can take it during the gap and the admin dies on bind, taking the
worker with it.

Keep the listener from the availability check and hand it to the admin.

* mini: clear the admin gRPC reservation before retaking it

A rerun inside one process would otherwise inherit the closed listener of
the previous run whenever the reservation fails, and the admin would accept
it and only find out inside Serve.

* mini: snapshot the admin options for the startup goroutine

The cleanup path read the package-level options long after the goroutine
started, so a later in-process run could have its reserved listener closed
by the previous run.
2026-08-24 14:55:53 -07:00
Chris Lu 51eb5333d3 ec: read a needle's intervals in parallel (#10911)
* ec: read a needle's intervals in parallel

A needle spanning more than one EC block gets one interval per block, and
consecutive blocks live on different shards. We read those intervals in
sequence, so a 4MB chunk landing in a volume's 1MB small-block region cost
five round trips to five different servers.

Read them concurrently into disjoint slices of a single buffer, at most 8 in
flight. Same change in the Rust volume server's phase C.

* ec test: seed the random payload instead of the deprecated rand.Read
2026-08-24 14:03:44 -07:00
Chris Lu 69cc2869ad Fixes from the review of the admin bucket policy UI (#10907)
* admin: treat a missing S3 Tables policy as an empty load, not an error

The bucket/table policy GET relayed the backend's 404 NoSuchPolicy to the
dialog, whose loader treats any non-OK response as a load failure and
keeps Save and Delete blocked. A bucket or table without a policy could
never be given one. Return policy null instead, the same contract
ShowBucketPolicy uses for classic buckets.

* admin: reject policy documents the structured editor would misread

A top-level JSON array passed the object guard (typeof [] is 'object')
and loaded as a zero-statement policy, which the next commit would
rewrite to an empty document. Object elements in Action/Resource were
coerced to '[object Object]' and saved that way on the s3tables surface,
which stores policies verbatim. Both now throw, which routes the
document to the JSON tab like other unrepresentable shapes.

* admin: let the JSON tab save documents the structured editor can't model

Save with the JSON tab active required a round-trip through
policyDocToEditorState, so exactly the documents the dialogs shunt to
'JSON tab only' mode (unrepresentable Effect, Resource+NotResource, and
the like) could never be saved - Delete was the only mutation left.
Invalid JSON still blocks; an unrepresentable document now saves and the
editor state stays marked unparsed.

* admin: pin the policy editor to what each consumer's backend supports

The s3tables evaluator has no NotResource/NotPrincipal fields - it
silently drops them, turning Allow+NotResource into allow-everything and
making Deny+NotPrincipal inert - and it only matches s3tables: actions
against s3tables ARNs, while the editor suggested s3: actions and
arn:aws:s3::: resources. New registerPolicyEditor knobs: allowNegation
hides the Not* modes and routes documents using them to the JSON tab;
resourceSuggestions pins the Resource autocomplete to the open
resource's ARN; the S3 Tables dialogs get an s3tables-only action
datalist. requirePrincipal now also hides NotPrincipal, which
policy_engine.ValidateBucketPolicy always rejects, and the client-side
check requires Principal specifically to match that server rule.

* admin: save S3 Tables policies from a button, not form submission

The multi-input structured editor sits inside a form whose Save button
was type=submit, so Enter in any single-line editor input - accepting an
autocomplete suggestion, say - implicitly submitted whatever half-built
statement the editor held, and the backend stores the document verbatim.
A lone statement with no Principal matches nobody, locking out every
non-owner. Save is now an ordinary button and the form ignores
submission.

* admin: block zero-statement policy saves

Committing the active tab before the emptiness check made 'Policy JSON
is required' dead code: an empty editor serializes to {"Statement":[]},
which the s3tables backend stores verbatim - evaluated default-deny for
every non-owner, while the statement-count column keeps showing 'Not
configured'. All three policy dialogs now refuse a save with no
statements and point at Delete instead. The classic bucket modal only
gained a clearer message; the server already rejected the document.

* admin: guard S3 Tables policy mutations against stale and overlapping requests

The save/delete completions ran against whatever resource the shared
modal happened to show by then: a slow PUT for one bucket would hide the
modal mid-edit of another and misattribute its alerts, a late DELETE
cleared the shared textarea over the newly opened resource with its
loaded flag set, and nothing stopped a double-click from firing two
overlapping mutations. Ported the classic modal's pattern: capture the
target on start, flag the mutation in flight with the buttons disabled,
and only touch the UI when the completion still matches the open
resource. Success now reloads the page, which also keeps the Policy
column's statement count honest.

* admin: confirm before deleting an S3 Tables policy

Delete Policy sat next to Save and fired on a single click; with
default-allow enabled one stray click silently dropped the resource
policy and left the bucket open to every principal. Same confirmation
the classic bucket modal already has.

* admin: let a corrupt stored bucket policy be shown, fixed, and deleted

A stored document the decoder rejects made the policy GET 500, and with
the loaded flag never set the modal blocked both Save and Delete - the
one policy an operator most needs to remove was the one they couldn't,
even though the delete path never reads the document. The GET now
returns the raw bytes alongside a null policy; the dialog hands them to
the JSON tab and unblocks the buttons.

* admin: url-encode the bucket name in the policy API calls

The filer lists any directory under the buckets path, names S3 would
never allow included; one carrying '#' or '%' broke the fetch URL or
addressed a different name than the modal shows.

* admin: drop stale edit-policy responses on the IAM policies page

The same race the bucket and S3 Tables dialogs already guard against:
open one policy's editor while its GET stalls, open another, and the
late response populates the editor under the second policy's name -
Update then saves the first policy's statements over the second.

* admin: warn before a bucket policy save drops unsupported fields

The editor tracks unmodeled top-level keys precisely so
confirmPolicyFieldDiscard can warn before the server's Version+Statement
decode discards them, but only the IAM page called it; the bucket modal
saved a pasted document with e.g. a console-generated Id without a word
while the editor kept displaying the field.

* s3: enforce the bucket policy size cap on both surfaces

The 20KB cap lived only in the admin UI, so a larger policy stored via
the S3 API displayed there but could never be re-saved, desyncing the
two writers the cap comment claimed could not desync. The constant now
lives in policy_engine next to the shared validator and PutBucketPolicy
rejects oversized documents with PolicyTooLarge, matching AWS.

* admin: ship the policy editor's fieldset styles with the editor

The .policy-stmt-* rules that undo Bootstrap's full-width legend reset
stayed behind in policies.templ when the editor markup moved to the
shared script, so the bucket and S3 Tables dialogs rendered Actions/
Resource/Principal as full-width jumbo headings. PolicyDatalists is the
component every consumer already renders once; the styles live there
now.

* s3: mirror bucket policy changes into the IAM store from the metadata subscription

The advanced-IAM path appends the bucket-policy:<bucket> document to
every STS/session evaluation, but only this gateway's own PutBucketPolicy
maintained that mirror - a policy tightened or created through the admin
UI (or another gateway) never reached it, so revoked access stayed live
indefinitely, and the delete side was an unimplemented TODO in any case.
The metadata subscription now diffs the stored policy on every bucket
entry change and updates or removes the mirror, covering all writers and
deletion with one mechanism; IAMManager gains the missing
RemoveBucketPolicy.

* admin: deduplicate the bucket policy write path

Set and Delete carried line-for-line identical filer closures;
bucketPolicyMutation already treats nil as clear-the-key. The shared
helper sits below Set's validation, since ValidatePolicy cannot take the
nil document Delete passes.

* s3: drop ValidateBucketPolicy's re-checks of ValidatePolicy rules

Both callers run ValidatePolicy first, which already enforces the
version and at-least-one-statement rules; the duplicates were dead code
with drifted error text.

* admin: seed a new statement's Resource from the pinned suggestions

A fresh statement on the S3 Tables dialogs started with no resource row
at all; seed it with the broadest pinned ARN the same way cfg.bucket
already seeds the classic modal.

* admin: refuse to save Not* fields the backend would silently drop

Hiding the NotResource/NotPrincipal modes was not enough where negation
is disallowed: the JSON tab accepts any valid document (that is its
job), and a statement's Advanced-fields box can reintroduce the keys, so
an s3tables save could still store fields the evaluator drops - turning
Allow+NotResource into allow-everything. commitPolicyActiveTab now runs
a final document-level check over what would actually be saved; Delete
stays available for cleanup.

* s3: move the IAM bucket policy mirror on a bucket rename

A same-directory rename delivers one event carrying both entries, and
the byte-equality short-circuit skipped the new name's mirror when the
policy was unchanged - while the replayed delete for the old name
removed its mirror, leaving the renamed bucket unmirrored. The mirror
decision is now a pure function that removes the old name and writes the
new one regardless of byte equality, with the rename cases unit tested.

* s3: backfill the IAM bucket policy mirror on lazy bucket loads

The metadata subscription only mirrors changes, so a policy that
predates the IAM integration never reached the bucket-policy:<bucket>
mirror and its grants did not bind on the IAM path until the policy was
next modified. The gateway is deliberately lazy at startup (nothing
lists all buckets), so the backfill hooks the same place a bucket's
policy first becomes known: the cold bucket-config load. EnsureBucketPolicy
writes only when no mirror is stored, so repeat loads cost one cached
read.

* s3: reconcile the bucket policy backfill against concurrent changes

The backfill's check-then-write could race an event-driven mirror update
or removal and re-store bytes that were already stale, with no later
event to heal it. EnsureBucketPolicy now reports whether it wrote, and a
write is reconciled against a fresh authoritative entry read: a changed
policy is re-mirrored, a removed one is removed. Anything changing after
that read fires its own event, which finds the backfill's write already
present and supersedes it. The backfill also carries the entry's raw
bytes rather than a re-marshaled document, so the reconcile can
byte-compare.

* s3: prime the bucket policy mirror before advanced-IAM authorization

The backfill ran from the lazy bucket-config load, but IAM authorization
evaluates the bucket-policy:<bucket> mirror before any handler runs - a
grant carried only by a not-yet-mirrored policy denied forever, and the
denied request never reached the code that would have loaded the bucket.
authorizeWithIAM now primes the bucket config first (an in-memory cache
hit once warm), and the backfill runs synchronously on the cold load so
the very first authorization already sees the mirror.
2026-08-24 00:52:01 -07:00
Chris Lu 68ec8ca655 admin: honor a persisted or admin.toml maintenance enabled=false (#10909)
* admin: honor a persisted or admin.toml maintenance enabled=false

The startup path discarded an operator's enabled=false twice over:
ApplyDefaultsToProtobuf treated the bool zero value as unset and applied
the schema default of true, and a force-enable migration block flipped
any survivor. With the legacy /maintenance UI routes gone, nothing could
write the config either, so the maintenance system ran unconditionally.

Keep the persisted enabled flag across schema-default application in
LoadMaintenanceConfig, drop the force-enable block, and add a top-level
[maintenance] enabled key to admin.toml as the config surface, persisted
through SaveMaintenanceConfig like the per-task settings. Absent config
still defaults to enabled.

* admin: track presence on the maintenance enabled flag

A plain proto3 bool cannot distinguish an operator's persisted false
from a legacy file that simply omits the field, so honoring false would
have silently switched maintenance off for configs written before the
toggle could be persisted. Make the field optional: files that predate
presence tracking keep the enabled default, while a file that explicitly
persists the toggle is honored either way.
2026-08-24 00:01:48 -07:00
Mathieu Arnold e931cccc7b Manage bucket policies via the admin ui (#10895)
* admin: manage S3 bucket policies from the admin UI

Bucket policies were only manageable through the S3 PutBucketPolicy API;
the admin UI had no equivalent to the quota/owner/lifecycle editors it
already offers. Add GET/PUT/DELETE for a bucket's policy, sharing the
exact validation the S3 gateway uses.

- Extract validateBucketPolicy/validateResourceForBucket out of
  s3api_bucket_policy_handlers.go into policy_engine.ValidateBucketPolicy /
  ResourceMatchesBucket so both the S3 API and the admin UI enforce
  identical rules.
- weed/admin/dash/bucket_policy.go: Get/Set/DeleteBucketPolicy, writing
  through ObjectTransaction + PATCH_EXTENDED (the lifecycle pattern) so a
  concurrent owner/quota/lifecycle change on the same bucket entry isn't
  clobbered. Propagation to every S3 gateway is automatic via the existing
  filer metadata log subscription. The S3 gateway's IAM policy mirror is
  deliberately not replicated here (its delete path is already an
  unimplemented TODO on the S3 side).
- New GET/PUT/DELETE /api/s3/buckets/{bucket}/policy routes, CSRF-guarded
  on writes.
- Bucket list and details modal now show a statement-count badge, read
  from the entry already fetched (no extra RPC).
- UI: a JSON-textarea policy editor modal, matching the lifecycle modal's
  structure.

* admin: reuse the visual policy editor for bucket policies

Extract the structured policy editor (add/remove statement, action/
resource/principal rows with autocomplete, JSON tab kept in sync) out of
policies.templ's inline script into a shared
weed/admin/static/js/policy_editor.js, and wire the bucket policy modal
in s3_buckets.templ up to it instead of a bare JSON textarea.

- registerPolicyEditor(which, config) replaces the hardcoded create/edit
  id derivation with a per-instance config (textarea/tab/body ids,
  datalist ids, requirePrincipal, bucket). The IAM policies page keeps its
  exact pre-extraction ids via two registerPolicyEditor calls, so its
  markup is unchanged.
- New policy_datalists.templ exposes the three shared <datalist>s
  (actions/resources/principals) as @PolicyDatalists(), now rendered by
  both policies.templ and s3_buckets.templ.
- requirePrincipal seeds new bucket-policy statements with Principal: "*"
  and adds a client-side check before save (the server, via
  policy_engine.ValidateBucketPolicy, remains the actual authority); the
  bucket config pins the Resource autocomplete to the open bucket instead
  of fetching every bucket in the cluster.
- layout.templ loads policy_editor.js globally, after admin.js/
  modal-alerts.js (basePath/escapeHtml/showAlert) which it depends on.

3a (the extraction) is a byte-preserving move verified against the
unchanged policies.templ behavior before layering 3b's parameterization
and the bucket-policy wiring on top.

* admin: migrate S3 Tables bucket/table policy editors to the shared editor

Third consumer of the shared visual policy editor: the S3 Tables bucket
and table policy modals (a bare JSON textarea each) now get the same
structured Editor/JSON tabs as the bucket policy and IAM policy pages,
via registerPolicyEditor('s3tablesBucketPolicy'/'s3tablesTablePolicy',
{ textareaId: ... }). Storage and validation are untouched - S3 Tables
policies still go through their own s3tables.PolicyDocument type and the
s3tables.policy extended attribute, unrelated to policy_engine and
s3-bucket-policy; only the editor UI is shared.

Fix a real bug surfaced by adding this second load path: the bucket
policy modal (and the naive first draft of this s3tables port) called
commitPolicyTextareaToEditor() right after a GET and then force-switched
to the Editor tab. commitPolicyTextareaToEditor() is designed to leave
the current tab in place and the editor state untouched when a document
fails to parse (so an in-progress edit survives a bad tab switch), so
forcing the Editor tab afterwards could show empty/stale editor state
that a careless Save would then serialize over a perfectly valid but
structurally-unusual stored policy. Add
loadPolicyTextareaIntoEditor(which) to policy_editor.js, which has no
"current tab" to defer to and instead falls back to the JSON tab with an
alert on a document the structured editor can't represent - the same
safety editPolicy already had in policies.templ - and use it at all three
"populate the editor right after a GET" call sites (bucket policy,
S3 Tables bucket policy, S3 Tables table policy).

* admin: show policy statement count on the S3 Tables buckets page

Mirrors the "Policy" column already added to the classic S3 buckets
list: a clickable badge with the statement count when the table bucket
has a resource policy, "Not configured" otherwise. S3 Tables policies
are a separate mechanism (s3tables.PolicyDocument under the
s3tables.policy extended attribute) from the S3 bucket policy work
elsewhere in this branch (policy_engine.PolicyDocument /
s3-bucket-policy), so this is a parallel implementation of the same
pattern rather than shared code.

- S3TablesBucketSummary gains PolicyStatementCount, populated in
  GetS3TablesBucketsData from entry.Entry.Extended[s3tables.ExtendedKeyPolicy]
  via the new extractS3TablesPolicyStatementCountFromEntry - no extra RPC,
  the entry is already fetched for ExtendedKeyMetadata.
- The badge reuses the existing .s3tables-bucket-policy-btn class, so it
  opens the same policy modal as the row's action button with no JS
  changes.

* admin: don't let a failed policy GET open the door to an empty overwrite

loadS3TablesBucketPolicy/loadS3TablesTablePolicy cleared the textarea,
then unconditionally called loadPolicyTextareaIntoEditor() regardless of
whether the GET actually succeeded - including when fetch() rejected or
the response was not ok, silently logged to console only. That leaves
the structured editor holding a legitimate-looking empty policy
({version, statements: []}), with the Editor tab active by default.

If Save is then clicked, commitPolicyActiveTab() serializes that empty
state into the textarea as `{"Version":"2012-10-17","Statement":[]}` -
a non-empty string - before the "Policy JSON is required" guard ever
sees it, so the guard passes and the transient load failure gets
written over whatever policy was actually stored.

Add s3tablesBucketPolicyLoaded/s3tablesTablePolicyLoaded, set true only
once a GET has actually completed (ok, including a genuinely empty
policy) and false on any failure path (fetch rejection or a non-ok
response, which previously fell through silently). Both submit handlers
now check the flag before touching the editor at all, and a failed load
surfaces via alert() instead of only a console.error - the user
previously had no visible indication the load had failed.

Verified with a jsdom simulation driving the real rendered page against
a stubbed fetch: a failed GET followed by Save now sends no PUT at all
(previously it sent Statement: []); a successful GET followed by Save
still PUTs the loaded policy unchanged.

* admin: address code review findings on the policy editor

1. policy_editor.js: policyEditors is only pre-populated for 'create'/
   'edit'; every other `which` (bucket, s3tablesBucket, s3tablesTable)
   stays undefined until its first successful async load. Nothing in
   this file enforces that a page hide its Editor/JSON tabs and
   Add-statement button until that load completes - the S3 Tables policy
   modals don't - so a click in that window (e.g. Add statement, or
   switching to the JSON tab) threw "Cannot read properties of undefined
   (reading 'unparsed')". Add policyEditorState(which), which lazily
   initializes a default state, and route addPolicyStatement, the
   jsonTabBtn 'show.bs.tab' handler, commitPolicyActiveTab, and
   renderPolicyEditor through it. Verified with a jsdom simulation
   against a never-resolving fetch: the exact click threw on the
   pre-fix code and no longer does.

2. s3_buckets.templ: the bucket-policy Save handler checked the
   textarea for emptiness before calling commitPolicyActiveTab(), which
   is what actually serializes the structured Editor tab's fields into
   that textarea. A policy entered entirely through the Editor tab (the
   primary path - never touching the JSON tab) left the textarea at
   whatever it was at load time, so creating a new policy this way hit
   "Enter a policy document" and Save silently did nothing. Move the
   commit before the emptiness check, preserving the existing alert and
   early-return. Verified with a jsdom simulation: Add-statement then
   Save (no tab switch) now PUTs the entered statement; before the fix
   the same sequence never reached fetch().

3. s3tables_buckets.templ / s3tables_tables.templ: the policy Editor/
   JSON nav-tabs were missing the ARIA roles Bootstrap's own tab pattern
   expects (role="tab"/"tabpanel", aria-selected, aria-controls,
   aria-labelledby) - screen readers had no way to tell these were tabs
   or which pane went with which button. Added the standard Bootstrap 5
   tab markup to both.

* admin: guard policy load/save flows against overlapping requests

1. s3tables.js: loadS3TablesBucketPolicy/loadS3TablesTablePolicy had no
   protection against overlapping loads. Opening one bucket's (or
   table's) policy dialog and then another's before the first GET
   resolved let the late response write its document into the shared
   textarea and mark the dialog "loaded" while it was now targeting the
   second resource - a subsequent Save would then push the first
   resource's policy onto the second. Add a per-load monotonic sequence
   number (s3tablesBucketPolicyRequestSeq / s3tablesTablePolicyRequestSeq,
   the same pattern already used for the classic bucket-policy load in
   s3_buckets.templ); a response is only applied - textarea, loaded flag,
   editor state - if its captured sequence still matches the latest one
   issued.

   Verified with a jsdom simulation: bucket A's policy load (artificially
   slow) followed immediately by bucket B's (fast) previously left A's
   policy in the textarea once A's late response landed; it now correctly
   keeps B's.

2. s3_buckets.templ: the bucket-policy Save button lives outside the
   (initially hidden) editor wrapper, so it stays clickable while a load
   is still in flight - the existing policyRequestSeq guard only protects
   the *load* from a stale response, not Save from firing before any
   load for the current bucket has completed. Add bucketPolicyLoaded,
   reset before each GET and set only once the matching response lands,
   and check it at the top of the Save handler.

   Verified with a jsdom simulation: clicking Save immediately after
   opening the dialog, before a (deliberately never-resolving) GET
   settles, now sends no PUT; a normal load-then-save sequence still
   PUTs the loaded policy unchanged.

* admin: address further code review findings on the policy editor

1. s3tables.js: loadS3TablesBucketPolicy/loadS3TablesTablePolicy only
   reset the JSON textarea when a new load starts; the structured editor
   kept showing the previously loaded resource's statements (Editor tab
   is the default active one) until the new fetch resolved. Call
   loadPolicyTextareaIntoEditor() against the now-cleared textarea
   immediately, so switching resources visibly resets the editor right
   away instead of only once its own load completes. Verified with jsdom:
   opening bucket A (loads fully) then bucket B (GET never resolves) no
   longer leaves A's statements visible in B's editor.

2. s3tables.js: deleteS3TablesBucketPolicy/deleteS3TablesTablePolicy had
   no loaded-state check, so a failed GET (which already blocks Save)
   left Delete fully able to remove the resource's stored policy sight
   unseen. Add the same s3tablesBucketPolicyLoaded/s3tablesTablePolicyLoaded
   guard Save already uses. Verified with jsdom: delete after a failed
   load now sends no DELETE; delete after a successful load is unaffected.

3. s3_buckets.templ: the bucket-policy Editor/JSON nav-tabs were missing
   the same ARIA roles already added to the S3 Tables policy tabs in an
   earlier round (role="tab"/"tabpanel", aria-selected, aria-controls,
   aria-labelledby) - this instance was out of scope for that review
   comment but is the same gap. Bootstrap's own tab.js already manages
   aria-selected on tab switch once the attribute exists, so no extra JS
   was needed.

4. s3_buckets.templ: neither the bucket-policy Save nor Delete handler
   guarded against a double-click, or against firing while the other was
   still in flight - two overlapping PUT/DELETE requests for the same
   bucket could land in either order. Add a shared
   bucketPolicyMutationInFlight flag: set (and both buttons disabled)
   before each fetch, cleared (and buttons re-enabled) on failure so the
   user can retry, left set through the existing success hide-and-reload
   path, and also reset when a new bucket's dialog opens so an abandoned
   in-flight request from a closed dialog can't leave the buttons stuck
   disabled. Verified with jsdom: double-clicking Save now sends exactly
   one PUT, and a Delete click while that PUT is still pending sends no
   DELETE.

* admin: scope bucket-policy mutation completions to the bucket that started them

1. The previous round's fix reset bucketPolicyMutationInFlight whenever a
   new bucket's policy dialog opened, to avoid leaving Save/Delete stuck
   disabled if the modal was closed mid-request. That traded one bug for
   a worse one: if bucket A's PUT/DELETE was still in flight when the
   user opened bucket B's dialog, the reset let B's Save/Delete fire
   immediately, and A's completion handler - unaware anything had
   changed - would still hide the (now B's) modal and reload the page
   out from under whatever the user was doing with B, on success, or
   alert a message with no bucket context, on failure.

   Stop resetting on reopen, so a pending mutation for a previous bucket
   keeps this bucket's Save/Delete blocked until it settles (matches the
   "preventing overlapping mutations" the review comment describes).
   Instead, capture policyEditorBucket as targetBucket right before each
   fetch and compare it against policyEditorBucket again in the
   completion handler: the in-flight flag is always released so the
   buttons never get stuck, but the modal-hide/reload/alert only fire if
   this bucket is still the one showing; a stale completion for an
   abandoned bucket just logs to the console instead.

   Verified with a jsdom simulation: opening bucket B while bucket A's
   Save is still pending leaves B's Save button disabled and a click on
   it a no-op; once A's PUT resolves, B's button re-enables but no
   modal.hide()/reload() fires (previously both fired unconditionally).

2. bucketPolicyDeleteBtn had no bucketPolicyLoaded check, unlike Save -
   a failed GET blocked Save but left Delete free to remove a policy the
   client never actually saw (the same gap already fixed for the S3
   Tables policy modals in an earlier round). Added the same guard,
   ahead of the confirm() dialog. Verified with jsdom: Delete after a
   failed load now sends no DELETE request.

* admin: fix spelling mistake
2026-08-23 22:11:18 -07:00
孙超 c80664ec21 s3: propagate storage rule fsync to volume server uploads (#10906)
The storage rule's fsync decision was computed by the filer
(detectStorageOption -> rule.Fsync) and applied on the filer's own HTTP
write path, but was never carried onto the chunk uploads S3 issues: the
AssignVolumeResponse had no fsync field, so the s3api client could not
learn the decision, and the chunked upload URL was hardcoded without it.
Every S3 write to a path with fsync configured went to the volume server
as a non-fsync write.

Carry the decision through the assign response:

- filer.proto: AssignVolumeResponse gains bool fsync, filled from the
  storage option the assign resolved.
- operation.AssignResult gains Fsync, so uploadChunk can append
  ?fsync=true to the volume server upload URL (single and replica
  fan-out paths).
- The S3 PUT/UploadPart assignFunc, the S3 copy path, the admin file
  browser upload, and the Iceberg worker assign functions all forward
  the response field.

Adds TestUploadReaderInChunksAppendsFsyncWhenAssigned.
2026-08-23 22:11:08 -07:00
Chris Lu 9c8d3b6a81 ec: refund the cleared leftover shards' slots in the encode source health check (#10903)
* erasure_coding: one home for the shard-count to volume-slots conversion

* ec: refund the cleared leftover shards' slots in the encode source health check
2026-08-23 21:48:20 -07:00
Chris Lu 36c97344ef s3: confine a Lance catalog table location to the caller's own bucket (#10901)
The Lance namespace gateway took the request-body location field, trimmed a
trailing slash, and passed it straight to the marker sink. That location feeds
TableDataDirFromMetadataLocation, which joins it under /buckets and collapses
any ../ segments, and writeMarker's CreateEntry then auto-creates every missing
parent. A caller could point the location at another tenant's bucket, or escape
/buckets entirely, and plant a fixed-name marker (recursively creating the
parents) or hide a victim's live table with .lance-deregistered.

Confine the declared location the way the Iceberg gateway already does: require
an s3:// URI whose bucket is the caller's own and whose path carries no
traversal segment, on both the declare and register handlers.
2026-08-23 11:49:52 -07:00
Chris Lu 74038e1b14 master: don't let a dead KeepConnected handler close its successor's channel (#10900)
A client that reconnects before the old handler exits re-registers the
same client name, and addClient overwrites the map entry. The old
handler's deferred deleteClient then closed whatever channel the map
held under that name: the new, live stream's. Receiving from a closed
channel returns nil immediately and forever, so the new handler's send
loop degenerated into sending empty responses at wire speed, pinning a
core on each side until the client killed the connection.

deleteClient now closes the channel its own handler registered and
leaves the map entry alone unless it still points to that channel. This
also closes the previously orphaned old channel, whose drain goroutine
used to leak. The send loop treats a closed channel as an exit instead
of a message stream.
2026-08-23 11:36:00 -07:00
Chris Lu cf0dba334c s3api: no filer failover after the callback has consumed part of a response (#10902)
s3api: no filer failover after fn has consumed part of a response

withFilerClientFailover replays fn verbatim on the next filer, so a filer
that died mid-stream followed by a healthy peer returned success with the
callback's closure-captured accumulator holding the dead filer's prefix
twice; the per-attempt accumulator in listWithRetry could not close this,
because the replay happens inside a single attempt. Track delivery on the
connection handed to fn: once a unary reply or streamed message has reached
the callback, surface the transport error unwrapped instead of failing
over, and let callers replay from a clean slate. A filer that fails before
delivering anything fails over exactly as before.
2026-08-23 11:30:43 -07:00
Chris Lu 0f85d005ad server: 416 only when no requested range overlaps, with Content-Range, and the Rust mirror (#10889)
* filer, volume server: return 416 when no requested range overlaps the content

* seaweed-volume: return 416 when no requested range overlaps the content

* server: check the range test error, use the request context, fix the no-overlap comment boundary
2026-08-23 11:13:36 -07:00
Chris Lu 173adbc291 master: never re-seed a raft cluster over committed state under -raftBootstrap (#10883)
* master: never re-seed a raft cluster over committed state

-raftBootstrap deleted logs.dat, stable.dat and snapshots on every start and
then bootstrapped a fresh cluster. Since hashicorp raft only snapshots after
8192 log entries, the TopologyId lives in the log, not in a snapshot, so the
pre-wipe snapshot recovery found nothing and each restart minted a new cluster
identity. A master that came up while it could not reach its peers seeded a
rival cluster; when the two logs met, SetTopologyId's split-brain guard fatally
stopped every master holding the other id, and the master layer crash-looped
with no quorum.

Bootstrapping is genesis. Drop the wipe and the inline bootstrap. The first
master in -peers already mints a cluster once it has confirmed no peer has a
leader, so the flag has nothing left to do and is now ignored; keeping that one
master the sole bootstrap authority is what stops a partition from minting two
clusters, so the flag must not widen it either. A master with state rejoins its
peers, and one whose data dir was reset is admitted by the sitting leader
instead of forking again.

* test: cover -raftBootstrap restarts in the multi-master suite

Three masters start with -raftBootstrap, the way the helm chart renders it on
every master on every roll, and the cluster has to hold one TopologyId after
they all restart. /dir/status is proxied to the leader, so each master's own
view of the identity is read out of its log, which is where a fork shows up.
Before the fix the hashicorp case minted a new id on each restart.
2026-08-23 11:10:20 -07:00
Junker der Provinz fa3bd5b5a7 mount: use the kernel-resolved node id in Link, not the persisted attribute (#10885)
* fix(mount): reply to LINK with the kernel node id, not the stored inode

Link() answered the kernel with out.NodeId = oldEntry.Attributes.Inode.
That attribute is a mount-runtime number and only entries created through a
mount carry one. An entry written by the S3 API, WebDAV or a direct filer
call persists inode 0, so the LINK reply named node id 0, which the kernel
rejects as invalid_nodeid and reports as EIO. The hard link itself had
already been written to the filer, which is why it looked correct again
after a mount restart.

The same stale number was also used as an inodeToPath key. AddPath(0, path)
filed the new link under inode 0, so a later Lookup on that name handed the
kernel node id 0 as well, and a LOOKUP reply carrying node id 0 means no
such entry.

in.Oldnodeid is the node id the kernel already holds for the source, and it
is the key inodeToPath is indexed by, so use it for the reply, for AddPath
and for the sibling sync.

Fixes #8404

* test(mount): cover the sibling sync in Link with a third hard link

The two existing cases never reach the body of syncHardLinkSiblings: with
two links the source alias and the name just created are both in skipPaths,
so the loop iterates over nothing and a change to that site goes unnoticed.
A third link leaves one name that no other part of Link() writes.

The new case drives three links off one source. It guards against covering
nothing (it fails if every path turns out to be a skipPath), checks that
every name of the file reports nlink 3, and then drives the sync with both
candidate keys to pin down which one it has to be: keyed by the source's
persisted Attributes.Inode, which is 0 for an entry written outside a mount,
GetAllPaths has no path to walk, while the kernel node id reaches the
sibling.

That second half is driven directly because Link() alone cannot tell the two
keys apart. The meta cache keeps one blob per hard link id (FilerStoreWrapper
setHardLink/maybeReadHardLink), so a read of any sibling returns the
attributes of the last write to any of them whether or not the sync ran.
2026-08-23 10:43:22 -07:00
Junker der Provinz 5ebc9c9f4b server: reject a Range start offset equal to the file size (#10898) 2026-08-23 08:20:25 -07:00
Chris Lu 3b8931c2f6 admin: address review feedback on the maintenance scanner fix 2026-08-23 00:26:38 -07:00
Chris Lu 8d8a25b1cf s3api: remove the duplicated listing retry helpers left by overlapping merges 2026-08-23 00:26:02 -07:00
Chris LuandJunker der Provinz c58795354a s3api: retry a transient filer failure on metadata listings (#10890)
* s3api: retry a transient failure when listing multipart uploads/parts

A blip on the way to the filer failed the whole ListMultipartUploads or
ListParts request. Both reported failure points sit inside one streaming
listing: the ListEntries call that opens the stream, and the stream.Recv
calls that drain it. Neither retried, so a single Unavailable answer from
a filer that was restarting turned into a 500 for the S3 client.

Replay the listing instead, bounded to three attempts with a 100ms
backoff that doubles. Only a transient failure is replayed. A not-found
answer stays authoritative so the empty-list branch still works, and
every other error still reaches the client on the first attempt.

This is scoped to (*S3ApiServer).list rather than added inside
DoSeaweedListWithSnapshot, which mount, the shell and the other object
listings share, and where a retry after a partial stream would
re-deliver entries the callback had already seen. Within one call to
list, a replay is safe: it collects into a fresh slice each time, so it
can neither duplicate nor drop entries.

That guarantee does not extend past this function. withFilerClientFailover
already re-runs its callback against the next filer on any non-NotFound
error without resetting the caller's accumulator, so on a multi-filer
gateway a mid-listing failover can itself produce a duplicated result
with err == nil, independent of this change and not fixed by it. Noted
in the PR rather than silently left for someone to rediscover.

Fixes #7221
References #7235

* s3api: move the listing retry inside list itself

---------

Co-authored-by: Junker der Provinz <jdp@braethoria.com>
2026-08-22 23:42:33 -07:00
Junker der Provinz 8d2c0273bd admin: stop the maintenance scanner pinning itself to one scan per second after a transient failure (#10887)
* admin: honour persisted task configs when building the maintenance policy

buildPolicyFromTaskConfigs passed a literal nil to vacuum, erasure_coding
and balance LoadConfigFromPersistence. Those functions look for their
LoadXTaskPolicy() accessor via a type assertion, which a nil interface can
never satisfy, so every call fell through to NewDefaultConfig() and the
policy came back with the compiled-in defaults - Enabled: true among them.
A task disabled on disk was therefore still scheduled, and the only trace
was a glog.V(1) "Using default ... configuration" line.

Thread the real ConfigPersistence through instead. There are two copies of
this function: the one in weed/admin/dash builds config.Policy on the
normal admin startup path and can simply take cp as its receiver, and the
one in weed/admin/maintenance is the fallback used when the config carries
no policy yet, which now receives the store from NewMaintenanceManager.
weed/admin/dash already imports weed/admin/maintenance, so the maintenance
side has to keep the duck-typed interface{} parameter that the task
loaders already use rather than importing the concrete type back.

The store is only handed over when a data directory is configured: an
unconfigured one has nothing to read, and a typed nil pointer would pass
the loaders' type assertion and then panic on first use.

Fixes #10874

* admin: restore the maintenance scan cadence after an error backoff

scanLoop shortens its ticker to the error backoff delay after a failed
scan, but it decided whether to replace the ticker by comparing the
target interval against the configured scan interval instead of against
the interval the ticker was actually running at. Once the errors stopped,
getScanInterval returned the configured interval again, the comparison
came out false, and the ticker was left at the backoff delay - so a
single transient scan failure pinned the scanner to one scan per second
for the rest of the process lifetime. That is the ~1/second cadence in
issue #10874: 658 KB/s of "Cancelled N stale pending balance tasks
before re-detection" and 193k orphaned task files over two days.

Track the interval the ticker is running at and compare against that, so
both entering the backoff and returning to the normal cadence replace the
ticker.

While in here:

- defer ticker.Stop() bound the ticker that was current when the defer
  was registered, so every replacement ticker leaked on return. Wrap it
  in a closure.
- running was written by Start/Stop and read by all three background
  loops without synchronisation. Guard it with the existing mutex, fold
  the running check in triggerScanInternal into the lock it already
  takes, and make Stop a no-op when not running so a second call cannot
  close the stop channel twice.

Refs #10874

* admin: make the maintenance policy actually reach the task detectors

Loading the persisted task configs into the maintenance policy only
matters if something reads that policy, and nothing did.

MaintenanceIntegration pushes the policy into every registered detector
and scheduler through interface{ SetEnabled(bool) } and
interface{ SetMaxConcurrent(int) } type assertions. Every task registered
through base.RegisterTask is backed by base.GenericDetector and
base.GenericScheduler, and neither implemented either method, so all four
assertions failed silently for every task on every startup. The policy's
enabled flag reached nothing: ScanWithTaskDetectors gates on
detector.IsEnabled(), and the queue's policy lookups for max concurrent
and repeat interval are fallbacks that only fire when the scheduler
reports zero, which the generic scheduler never does.

Add the setters, delegating to the TaskConfig.SetEnabled the interface
already declares and to TaskDefinition.MaxConcurrent, which is what
GetMaxConcurrent returns.

Applying the policy required three more fixes, because with the
assertions working the policy could now do damage as well as good:

- IsTaskEnabled reports false for a task type the policy has no entry
  for, so applying it unconditionally would have disabled every task the
  policy does not list. Skip task types with no policy entry: no entry
  means no opinion, not disabled.

- ec_balance was exactly such a task. It is registered like the other
  three but had no entry in the policy builder and no accessor on
  ConfigPersistence at all, so its configuration could never be
  persisted. Add SaveEcBalanceTaskPolicy/LoadEcBalanceTaskPolicy, the
  task_ec_balance.pb file, the SaveTaskPolicy dispatcher case, and the
  policy entry.

- InitMaintenanceManager ran before loadTaskConfigurationsFromPersistence,
  which replaces each task's whole config object, so the policy was
  applied and then immediately thrown away. Swap the order. Both read the
  same files, so the policy is now the last writer and stays
  authoritative.

MaintenanceManager.UpdateConfig also updated the queue's and the
scanner's policy but not the integration's, so a policy changed at
runtime never reached the detectors. Add MaintenanceIntegration.SetPolicy
and call it.

While building the policy, stop hand-copying each task's fields and use
the task's own ToTaskPolicy(). The hand-written version was a second
definition of every task's policy and had already lost the erasure coding
preferred tags and replica placement and the balance IO rate limit. For
the same reason, the "nothing persisted yet" branches of
LoadVacuumTaskPolicy, LoadErasureCodingTaskPolicy and
LoadBalanceTaskPolicy now derive from each task's NewDefaultConfig()
instead of a third hand-written copy. Those copies had drifted, so with a
data directory but no config file on disk the effective defaults differed
from what the task and the admin UI schema both advertise:

  vacuum          scan interval  24h  -> 2h
  balance         scan interval   6h  -> 30m
  balance         imbalance      0.1  -> 0.2
  erasure coding  scan interval 168h  -> 1h
  erasure coding  fullness      0.90  -> 0.95
  erasure coding  min volume   1024MB -> 30MB

Finally, weed/admin/dash and weed/admin/maintenance each carried a copy
of the policy builder and they had already diverged. Export the
maintenance one as BuildPolicyFromTaskConfigs and have dash call it.

Refs #10874

* worker: warn when a config store cannot supply a task's persisted config

LoadConfigFromPersistence logged a single glog.V(1) "Using default X
configuration" for every way of not loading anything, so the bug in
issue #10874 - a store handed in that the type assertion rejects, leaving
a task running on compiled-in defaults - looked exactly like the normal
"no data directory configured" case. The reporter had to read the source
to work out why their disabled task kept running, and asked for this
specifically.

Separate the cases. A non-nil store that does not provide the accessor is
always a wiring bug and is now logged at warning level, naming the type
and the missing method. A read error or a policy that will not apply is
also a warning. No persistence configured, and a store with nothing saved
yet, stay at V(1): those are normal.

Refs #10874

* admin: stop GetTaskPolicy panicking on a maintenance policy that is nil

GetTaskPolicy dereferenced its MaintenancePolicy argument to look at
TaskPolicies, so IsTaskEnabled, GetMaxConcurrent and GetRepeatInterval
all took the admin process down when handed a nil policy. A nil policy is
not a programming error here: MaintenanceConfig.Policy is unset until
something builds one, DefaultMaintenanceConfig returns a config with no
policy at all, and UpdateConfig installs whatever config it is given.
Found by calling IsTaskEnabled with the policy from a freshly defaulted
MaintenanceConfig.

Treat a nil policy as "no entry": no task enabled, the safe concurrency
default of 1, and a repeat interval of 0 so callers fall back to their
own default instead of reading DefaultRepeatIntervalSeconds off nil.

Also add the startup test this was found with. It walks the admin
server's startup sequence over a data directory that has balance saved as
disabled and checks the state that decides whether issue #10874 happens:
the balance detector reports disabled, vacuum stays enabled, and tasks
whose config was never saved keep their compiled-in default.

Refs #10874

* admin: document the synchronisation SetPolicy would need beyond startup

ConfigureTasksFromPolicy now really writes TaskDefinition.Config and
TaskDefinition.MaxConcurrent, which the scan loop reads through
detector.IsEnabled() with nothing synchronising the two. Every caller
runs during admin server startup today, before the scan loop exists, so
there is no live race - but the next caller has to add the locking, and
the same already applies to UpdateAllConfigs replacing the whole config
object. Write it down at the seam instead of leaving it to be
rediscovered.

Refs #10874
2026-08-22 23:41:45 -07:00
Junker der Provinz f710b6003a s3api: retry a transient failure when listing multipart uploads/parts (RFC on layering) (#10886)
s3api: retry a transient failure when listing multipart uploads/parts

A blip on the way to the filer failed the whole ListMultipartUploads or
ListParts request. Both reported failure points sit inside one streaming
listing: the ListEntries call that opens the stream, and the stream.Recv
calls that drain it. Neither retried, so a single Unavailable answer from
a filer that was restarting turned into a 500 for the S3 client.

Replay the listing instead, bounded to three attempts with a 100ms
backoff that doubles. Only a transient failure is replayed. A not-found
answer stays authoritative so the empty-list branch still works, and
every other error still reaches the client on the first attempt.

This is scoped to (*S3ApiServer).list rather than added inside
DoSeaweedListWithSnapshot, which mount, the shell and the other object
listings share, and where a retry after a partial stream would
re-deliver entries the callback had already seen. Within one call to
list, a replay is safe: it collects into a fresh slice each time, so it
can neither duplicate nor drop entries.

That guarantee does not extend past this function. withFilerClientFailover
already re-runs its callback against the next filer on any non-NotFound
error without resetting the caller's accumulator, so on a multi-filer
gateway a mid-listing failover can itself produce a duplicated result
with err == nil, independent of this change and not fixed by it. Noted
in the PR rather than silently left for someone to rediscover.

Fixes #7221
References #7235
2026-08-22 23:00:50 -07:00
Junker der Provinz f3caf6e7da admin: count plugin-runtime workers in worker metrics (#10884)
* admin: count plugin-runtime workers in worker metrics

The admin server keeps two worker registries: the legacy maintenance-worker
map, filled by workers registering over the worker gRPC stream, and the plugin
worker registry, filled by workers started as `weed worker`. Both the
SeaweedFS_admin_workers_connected / SeaweedFS_admin_worker_slots gauges and the
dashboard's Workers card read only the legacy map, so a cluster that runs the
admin and its workers as separate components reported 0 workers even while its
workers showed up on the plugin pages and ran scheduled jobs.

Aggregate both registries instead. The two are merged by worker ID: `weed mini`
starts both runtimes out of one working directory, so they share the persisted
worker ID and must not be counted twice. For such a worker the slot numbers
still come from the legacy registry, which keeps mini's existing readings.
Plugin workers report their slots in the heartbeat, so detection and execution
slots are summed from there; a worker that has connected but not yet sent a
heartbeat counts as connected with zero slots.

Fixes #10525

* admin: clamp negative worker-reported slot values in metrics merge

A plugin worker's self-reported heartbeat slot counts are untrusted
input; clamp them to 0 before summing so a stale or misbehaving
worker can't drive the aggregate gauge negative, matching the same
defensiveness already used in registry.go's own slot arithmetic.
2026-08-22 22:54:39 -07:00
641fc8b031 admin: add visual iam policy editor (#10878)
* admin: add visual iam policy editor

Add a structured, tabbed editor (Editor / JSON) for creating and editing
IAM policies in the admin dashboard, alongside the existing raw-JSON
textarea:

- policies.templ: per-statement cards for Sid, Effect, Action, and
  Resource, with unmanaged fields (Principal, NotPrincipal, NotResource,
  Condition, or anything else) preserved verbatim in a per-statement
  "advanced fields" JSON box so nothing is lost on round-trip. Switching
  tabs commits and reparses in both directions. Restored the "Use Sample
  Policy" button, now filling both the structured editor and the JSON
  tab. The "Validate" button now calls the existing but previously
  unused POST /api/object-store/policies/validate endpoint instead of
  doing JS-only checks.
- Progressive Resource ARN autocomplete: suggests bucket names first,
  then once "bucket/" is typed, suggests bucket/* plus the bucket's
  direct subfolders, drilling down one path segment at a time as the
  user types further "/" characters.
- New GET /api/files/list-folders endpoint (file_browser_handlers.go)
  backing the folder autocomplete: wraps the existing file browser data
  function and returns just the subdirectory names as JSON, scoped to
  paths under /buckets.
- Action-name suggestions (datalist) for the Action field, sourced from
  the existing s3_constants.S3_ACTION_* constants plus new
  s3_constants.S3TABLES_ACTION_* constants (extracted from the s3tables
  operation dispatch switch) so the suggestion list can't drift from the
  strings the engines actually understand.
- policy_handlers.go: ValidatePolicy now accepts a statement with only
  NotResource set (previously required Resource), matching
  policy_engine.validateStatement and the fact the new editor makes such
  statements reachable from the UI.
- Tests: ValidatePolicy behavior, route registration for the policy API
  and the new list-folders endpoint, list-folders path scoping, and the
  action-suggestion list's shape.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: fix XSS, cache poisoning, and cap overshoot in policy editor

Address code review findings on the IAM policy editor added in the
previous commit:

- policies.templ (displayPolicyDetails): escape every interpolated
  policy value (Sid, Effect, Action, Resource, policy name, and the raw
  JSON document) before assigning to innerHTML. Policy documents can
  come from other admins or an import, so an unescaped field could
  execute script when the "View" modal renders it.
- policies.templ (policyEditorStateToDoc): reject JSON arrays in a
  statement's "advanced fields" box, not just invalid JSON. `typeof []
  === 'object'` was true, so a JSON array was assigned to the statement;
  subsequent property assignments (Sid, Effect, ...) landed on the array
  object but JSON.stringify of an array only serializes numeric indices,
  silently dropping them.
- policies.templ (loadPolicyFolderNames): on a failed folder lookup,
  remove the cache entry instead of permanently caching the empty
  fallback, so a transient network/server error doesn't block retries
  for the rest of the page's lifetime.
- file_browser_handlers.go (ListFolders): stop appending directory
  names as soon as the running count reaches maxListFoldersEntries,
  instead of only checking the cap after a full page is processed,
  so the returned list never exceeds the configured cap.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001
to keep the diff scoped to this file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: stop policy editor from clobbering the active tab and dropping malformed advanced fields

Address two review findings on the IAM policy editor (Issue 3, stored-XSS
in displayPolicyDetails, was already fixed by the previous commit and is
unchanged here):

- createPolicy, updatePolicy, and validatePolicyDocument always committed
  the structured editor's (possibly stale) state into the JSON textarea
  before submitting, even when the user had just edited the JSON tab
  directly. That silently discarded the user's JSON edits and
  validated/saved the old structured-editor state instead, which could
  leave broader permissions in force than intended.

  Added commitPolicyActiveTab(which), which commits whichever tab is
  currently visible into the other side instead of unconditionally
  overwriting the JSON tab from the editor: if the JSON tab is active it
  parses that JSON back into the structured editor (without touching the
  textarea itself), otherwise it serializes the structured editor into
  the textarea as before. All three call sites, plus the JSON-tab
  "show.bs.tab" handler, now use this and abort with an alert if the
  currently active tab's content can't be committed.

- policyEditorStateToDoc silently continued with an empty object when a
  statement's "advanced fields" box held invalid JSON, so switching
  tabs, validating, or saving would drop Principal/NotResource/Condition
  from that statement without telling the user. It now throws (with the
  statement number and parse error) on invalid or non-object JSON there,
  and callers surface that via showAlert and abort instead of proceeding.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: keep unmanaged top-level policy fields across editor tab switches

policyDocToEditorState only carried Version and Statement into editor
state, so any other top-level key (e.g. Id) present in the JSON tab was
silently rewritten away as soon as the user switched to the Editor tab
and back. Capture those keys in state.otherFields and merge them back in
policyEditorStateToDoc before Version and Statement are written, so the
two tabs stay faithful to each other and the editor never rewrites text
the user typed.

Note this is editor fidelity only: the admin API's
policy_engine.PolicyDocument carries just Version and Statement, and
DocumentJSON is never populated, so such fields are still discarded by
the server once a policy is saved. Making them survive a save would
require a backend change, which is out of scope here.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: warn before a policy save discards unsupported top-level fields

The editor round-trips unmanaged top-level keys (e.g. Id) between the
Editor and JSON tabs, but the admin API's policy_engine.PolicyDocument
carries only Version and Statement, so the server drops them on save and
the user saw no indication.

Added confirmPolicyFieldDiscard(), called from createPolicy and
updatePolicy after the active tab is committed (so the field list is
accurate whichever tab is showing). It names the fields that will be
lost and lets the user confirm or cancel. Not wired into
validatePolicyDocument, which doesn't persist anything.

Chose the warning over the alternative of persisting these fields
through the backend: policy_engine.PolicyDocument is shared by the S3
bucket-policy engine and IAM evaluation, so extending it would change
the stored document shape for every policy in the codebase - far beyond
the scope of this editor.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: reject malformed Effect and Resource/NotResource conflicts in policy editor

Two review findings on the IAM policy editor:

- policyDocToEditorState defaulted any non-"Deny" Effect (missing,
  misspelled, wrong case) to "Allow". A statement meant to be "Deny" with
  a typo like "deny" would silently become a permissive "Allow" instead
  of being rejected. It now throws on anything but an exact "Allow" or
  "Deny", naming the offending statement and value.
  commitPolicyTextareaToEditor catches this the same way it already
  catches invalid JSON: alert the user and keep the JSON tab active
  instead of switching to the Editor tab with wrong data.

- policyEditorStateToDoc could save a statement with both Resource (from
  the structured field) and NotResource (surviving in the "advanced
  fields" extras from before the user switched to using Resource) set at
  once - a contradictory combination neither the admin's ValidatePolicy
  handler nor policy_engine's evaluator rejected. When the structured
  Resource field is non-empty it now deletes any leftover NotResource
  from extras, consistent with the file's existing rule that structured
  fields take precedence over extras. Mirrored the existing
  Principal/NotPrincipal exclusivity check in
  weed/admin/handlers/policy_handlers.go's ValidatePolicy to reject the
  same combination server-side, since create/update perform no
  validation at all. Deliberately left policy_engine.validateStatement
  (used by the S3 bucket-policy PUT handler for every bucket policy in
  the product) unchanged - extending that shared validator is a larger,
  separate change outside this admin-editor fix's scope.

Added a handler test for the new Resource+NotResource rejection.
Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: add NotResource support to the visual policy editor

Since Resource and NotResource are mutually exclusive (enforced by a
previous fix), NotResource could previously only be set through the raw
JSON in a statement's "advanced fields" box. Promote it to a first-class
mode of the structured editor:

- The static "Resources" label is now a Resource/NotResource dropdown;
  the same list of values underneath is reused for either key depending
  on the selected mode, with a short form-text explaining the semantics.
- NotResource is added to POLICY_STATEMENT_KNOWN_KEYS, since it's now a
  managed field like Resource rather than something that falls through
  to extras.
- policyDocToEditorState derives resourceMode from which key is present
  on load, and throws (same handling as the existing malformed-Effect
  case: alert, keep the JSON tab active) if a hand-edited document has
  both Resource and NotResource on one statement, since that can't be
  represented by the dropdown.
- policyEditorStateToDoc writes only the key matching the selected mode,
  replacing the previous one-directional "delete NotResource whenever
  Resource is set" fix with mode-driven logic that also deletes Resource
  when NotResource is selected.
- displayPolicyDetails (the read-only View modal) now shows the actual
  NotResource values with a distinct label instead of a static
  "(NotResource used instead)" placeholder.

No backend changes: the server-side "cannot specify both" check added
previously in policy_handlers.go's ValidatePolicy already covers this.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: reject non-object policy documents; catch Principal/NotPrincipal conflicts server-side

Two review findings:

- policyDocToEditorState treated a top-level JSON value that wasn't an
  object (null, or a bare string/number/boolean) as an empty statement
  list instead of failing explicitly. If the user typed e.g. "hello" or
  42 in the JSON tab and switched to the Editor tab, their input was
  silently discarded and replaced with an empty policy - the same class
  of "guess instead of reject" bug fixed for malformed Effect and
  Resource/NotResource conflicts previously. Added an explicit check
  that throws for null/scalar input, while leaving array and object
  document shapes accepted exactly as before.

- weed/admin/handlers/policy_handlers.go's ValidatePolicy checked the
  Resource/NotResource conflict by non-empty length
  (len(...Strings()) > 0), which misses a statement where Resource is
  explicitly present but an empty list (e.g. "Resource": []) alongside a
  non-empty NotResource. Switched that check to field presence (!= nil),
  matching how policy_engine's own validateStatement already treats
  Principal/NotPrincipal exclusivity. Also added the equivalent
  Principal/NotPrincipal presence check to this handler, which had none
  before - the advanced-fields box in the visual editor lets a user set
  both today, and nothing server-side caught it. The existing
  non-empty "Resource or NotResource is required" check is left as a
  length check, since an empty array shouldn't count as "provided".

Added test cases for both conflict checks in policy_handlers_test.go.
Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: add Principal/NotPrincipal support to the visual policy editor (v1, AWS-only)

Adds a first, deliberately narrow structured editor for a statement's
Principal/NotPrincipal, left out when NotResource support was added:

- A Principal/NotPrincipal mode dropdown mirrors the existing
  Resource/NotResource one (same mutual-exclusivity handling: the two
  fields can't be set at once, and switching modes reuses the same
  value list).
- A simple repeatable text-value list feeds a single {"AWS": [...]}
  object on save - always the AWS type, never the "bare" (untyped)
  SeaweedFS-extension shape. Per policy_engine's allowedPrincipalKeys,
  Service/Federated/CanonicalUser also parse successfully, but nothing
  in the S3 bucket-policy evaluation path ever sets a real caller's
  principal to a service name, an OIDC provider ARN, or a canonical
  user ID, so only AWS is functionally meaningful today - out of scope
  for this v1.
- On load, only the exact {"AWS": ...} single-key shape is unwrapped
  into the structured field and removed from "extras". Anything else
  (bare string/array, a different single type key, or several type
  keys at once) is left untouched in "extras" exactly as before, with a
  visible warning under the dropdown so the user knows a
  Principal/NotPrincipal exists but isn't shown there. Saving with the
  structured field left empty never touches whatever's already in
  extras, so a preserved complex form isn't silently dropped just
  because the user didn't touch this field.
- The read-only View modal now displays Principal/NotPrincipal for any
  shape (via a small generic summarizer), not just the AWS-simple one.
- Generalized the action/resource field-to-state-key mapping (used by
  commitPolicyEditorForm and the add/remove-item click handler) into a
  shared lookup table instead of stacking another ternary, now that a
  third field (principal) exists.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: support the bare "*" wildcard Principal in the visual editor

"Principal": "*" (and NotPrincipal: "*") is the standard AWS shorthand
for "everyone" and is common in real bucket policies, but the v1
Principal/NotPrincipal editor only recognized the {"AWS": ...} object
form, leaving a bare "*" statement's principal hidden in Advanced
fields.

parseSimpleAwsPrincipal now also accepts the bare string "*" as a
simple, structurally-editable value. On save, a principal value list
containing exactly ["*"] is written back as the bare "*" string
(matching the common convention) rather than wrapped as {"AWS": "*"};
anything else still wraps under AWS as before. Updated the field's
form-text hint accordingly.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: add Principal field autocomplete backed by users + IAM roles

Adds a datalist-backed autocomplete for the policy editor's Principal/
NotPrincipal text fields, sourced from a new API listing existing
identities:

- weed/admin/dash/principal_suggestions.go: AdminServer.GetPrincipalSuggestions
  combines S3 user ARNs (via the existing GetObjectStoreUsers +
  iam.UserArn) with IAM role ARNs (via integration.NewFilerRoleStore /
  ListRoles, reusing the exact same construction already used in
  iam_manager.go - no new dependency risk introduced). Role ARNs are
  reconstructed from the role name using SeaweedFS's default
  arn:aws:iam::role/<name> convention rather than fetching each role's
  stored definition, since this only backs a suggestion list. Role
  listing failures are logged and swallowed rather than failing the
  whole request - an incomplete suggestion list is fine, blocking
  policy editing over it is not. Service accounts are deliberately not
  listed separately: a service account's ARN is identical to its parent
  user's, already covered by the user list.
- weed/admin/handlers/policy_handlers.go: GetPrincipalSuggestions handler
  exposing this as {"principals": [...]}.
- Route registered at the API root (GET /api/principals) rather than
  under policyApi's "/object-store/policies" prefix, since that
  subrouter's existing "/{name}" GET route would shadow any
  single-segment GET route registered after it (the same class of
  gotcha previously seen with "/validate").
- weed/admin/view/app/policies.templ: a shared, lazily-fetched-once
  policyPrincipalSuggestions datalist (flat list - unlike the
  progressive per-folder Resource ARN autocomplete, users/roles aren't
  hierarchical), wired into policyListRowHtml for field:"principal" and
  populated on input/focus, with "*" always offered first.

Added tests for the new ARN-construction helper and route registration.
Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: fix fieldset/legend styling in the structured policy editor

Bootstrap's form reset stretches <legend> to the fieldset's full width
(float: left; width: 100%), which loses the native "notch in the
border" look and makes each section's label bar as wide as the card.

Add two scoped classes: .policy-stmt-fieldset (border, rounded
corners, spacing between sections) and .policy-stmt-legend (undoes the
float/width so the legend hugs its content, with a little padding).
Applied to the three per-statement sections (Actions,
Resource/NotResource, Principal/NotPrincipal), replacing the ad hoc
"border rounded" utility classes that were doubling up with the
fieldset's own border. Also gave the "Advanced fields" <details> a
small top margin to match the new spacing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: suggest bucket/* alongside the bucket itself in Resource autocomplete

At the bucket-name stage of the Resource field's progressive
autocomplete, only "arn:aws:s3:::bucket" was offered. Add
"arn:aws:s3:::bucket/*" right alongside it, since granting access to
everything in a bucket is the more common case and previously required
typing a "/" first to reach the folder-level "*" suggestion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: keep an unparseable policy in the JSON tab instead of wiping it

editPolicy() built the structured state inside the fetch .then, so a
policy the editor cannot model threw into the sibling .catch, which
alerted and called hide(). Showing the alert at that moment left the
modal on screen with an empty editor and the document only in the JSON
tab, and Save Changes then serialized the empty state over the policy.

Reachable two ways, since neither create path rejects these: the admin
API never validates on create, so "Effect":"allow" is stored as typed,
and policy_engine.validateStatement lets Resource and NotResource sit
in the same statement.

Hand the document to the JSON tab instead, which is what that tab is
for, and mark the state so nothing serializes the placeholder over it.

* admin: validate a policy document before saving it

Validation was wired only to the Validate button, so nothing stopped a
document the server's own validator rejects from being stored. With the
structured editor supplying the boilerplate and required dropped from
the textarea, opening the modal, typing a name and clicking Create
Policy was enough to save a statement-less policy.

Share validatePolicyJSON with the two save paths and abort on failure.

* admin: bound the folder autocomplete listing

maxListFoldersEntries caps the folders collected, but nothing capped the
entries paged through to find them, so a bucket holding only flat object
keys - no subfolders to count - was walked to the end, 200 entries per
round trip, behind one keystroke. Measured against an in-process filer:
6 entries 0.5ms, 3k entries 7.7ms, 30k entries 53ms, all of it linear in
the directory rather than in the answer.

Cap the scan as well, and let GetFileBrowser take a prefix so the segment
the user is still typing is filtered by the filer instead of by paging.
The same 30k directory now answers in 0.6ms once a prefix is typed.

* admin: clean the path before scoping list-folders to /buckets

util.CleanWindowsPath only rewrites backslashes, so "/buckets/../etc"
walked straight past the prefix check the endpoint relies on for its
scope. Nothing leaked - filer paths are literal keys, so the traversal
resolved to nothing - but the check reads as a boundary and wasn't one,
and the test asserting it didn't cover the one input that would try.

validateAndCleanFilePath in the same file already does this.

* admin: stringify policy values before escaping them

escapeHtml calls text.replace directly, and the Sid, the per-item action
and resource inputs, and the View modal's Resource/NotResource all pass
values straight out of JSON.parse. A policy carrying "Sid": 5 or
"Action": [1] threw "text.replace is not a function" and took the render
with it. escapedJoin already coerced; use it everywhere and coerce the
editor state at the point it's built.

* admin: only show the NotResource hint in NotResource mode

The hint rendered unconditionally, so it sat under a selector reading
"Resource" telling the user the statement applies to everything except
what they'd listed. Redraw the card when the selector changes so it
follows the mode.

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-22 12:33:08 -07:00
github-actions[bot] 3563738699 4.44 2026-08-22 07:41:01 +00:00
Chris Lu c1a993bc3b filer: keep the TUS sub-chunks that already landed when a write fails (#10876)
* filer: keep the TUS sub-chunks that already landed when a write fails

A PATCH is split into 4MB sub-chunks, and each one is recorded in the
session as soon as it is stored. The session listing is what HEAD reports
as Upload-Offset and what the final entry is assembled from, so a record
is a promise that the data behind it exists.

When a later sub-chunk failed - a read-only volume, or a client that hung
up mid-body - the error path deleted the needles of every sub-chunk the
same PATCH had written but left their records in place. The resuming
client was then told to continue past bytes the filer had just queued for
deletion, and the upload completed into a gapless manifest pointing at
needles that were gone: HEAD returned the right size, GET died mid-body
once a vacuum reclaimed them.

Recorded sub-chunks now stay, which is what resumption expects: the
client picks up at the offset the session reports, and an upload that is
abandoned frees its chunks with the session.

* filer: drop a TUS chunk's record before freeing its data

filer.CreateEntry can return an error with the entry already inserted -
the parent-directory pass runs after the insert and keeps the entry when
it fails. A failed saveTusChunk therefore does not mean the record is
absent, and deleting the needle outright left the same corruption the
resume path used to cause: a session record pointing at data that is gone.

Remove the record first and only free the needle once it is gone. A
record lost with its data still stored merely leaks, which the vacuum and
fsck paths already account for.

* test: cover a TUS PATCH that is cut off mid-body

Resets the connection after one 4MB sub-chunk has landed, resumes from the
offset the session reports, and vacuums before reading the file back, so
anything the filer deleted behind a kept record shows up as a short read.
2026-08-22 00:30:14 -07:00
df93d01c06 admin: add bucket lifecycle rule editing (#10860)
* admin: add bucket lifecycle rule editing

* address greptile's comments

* more small fixes

* coderabbit's comments

* more comment fixes

* more fixes

* more

* maybe last

* last ?

* 14850

* 14851

* filer: stamp the content MD5 on every SaveInsideFiler write

An entry's ETag falls back to Attributes.Md5, so conditional writers key
IF_ETAG_MATCH off it. SaveInsideFiler carried the looked-up attributes
forward without refreshing the hash, leaving it describing whatever the
previous writer stored: a later conditional write matched the stale hash
and overwrote content that had already changed.

* s3api: give the bucket lifecycle constants and the write route key one definition each

The extended-attribute keys, the XML size cap and the object-write ring key
prefix were each spelled out in two places, so the admin dashboard's copies
could drift from the gateway's. Move them to the packages both sides already
import and alias them where the short local name reads better.

* admin: patch the bucket entry's lifecycle keys instead of rewriting the entry

The save read the bucket entry, edited its extended map and wrote the whole
entry back, guarded by IF_UNMODIFIED_SINCE. Nothing that writes a bucket
entry advances its mtime - not the S3 gateway's patchBucketEntry, not
SetBucketOwner, not SetBucketQuota - so the guard never fired and the stale
snapshot reverted whatever else had changed since the lookup.

Send the PATCH_EXTENDED mutation the S3 gateway already uses for these keys:
the filer re-reads and merges under the bucket path lock, so only the two
lifecycle keys move. That removes the reason for the mtime snapshot, the
verification retry loop and the compensating restore of the cleared day-TTL
rules, which the migration now logs instead.

* s3api: run the delete-lifecycle day-TTL migration through the shared helper

DeleteBucketLifecycleHandler kept its own copy of the read-strip-write
sequence the put handler now shares, including a missing return that let a
ToText failure persist a truncated filer.conf and write a second response.
It also wrote the whole file back unconditionally, reverting any concurrent
edit; the shared helper writes conditionally.

* admin: answer 404 when a lifecycle request names a bucket that does not exist

Every SetBucketLifecycle failure came back as 500, including the lookup miss
for an unknown bucket, so a client or monitor read a caller error as a server
fault and retried it.

* s3api: emit lifecycle XML a client would recognize

Two changes to what MarshalCanonical writes, both visible through
GetBucketLifecycleConfiguration, which replays the stored bytes verbatim:
stamp the S3 namespace on the root, and put a size range under <And>. A
<Filter> carries one predicate, so two size bounds side by side is a shape
AWS does not document. Parsing still accepts either.

* admin: fix the lifecycle editor's handling of stored status, deletes and empty saves

Four things the editor got wrong:

A stored <Status> the S3 API never validated, say 'enabled', left both radio
buttons unchecked, so reading the form threw on a null querySelector result
and Save did nothing. Collapse anything but an exact 'Enabled' to 'Disabled',
which is what the engine already does with it.

Deleting a rule re-rendered an open edit form from the snapshot taken when
editing began, discarding what had been typed; every other transition folds
the form in first.

The Transition warning only matched a bare <Transition>, missing the form
with attributes, self-closed or namespace-prefixed.

Saving an emptied rule list clears the configuration through a path with no
prompt, next to a Delete-all-rules button that asks.

Also collapses the three divergent copies of formatBytes on this page to one.

* filer: stop the day-TTL migration from deleting an operator's path rule

The migration removed every rule under the bucket's path that carried a day
TTL in the bucket's collection. The add path it is retiring used
AddLocationConf, which merged its TTL onto whatever already sat at the
prefix, so a rule can hold operator settings the lifecycle path never wrote -
a disk type, WORM retention, a read-only flag, a placement pin. Deleting the
whole rule to retire its TTL took those with it, leaving objects under that
prefix on defaults nobody asked for.

Delete only rules shaped like ones the add path created from scratch;
anything else keeps its settings and loses just the TTL.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-08-21 23:42:26 -07:00
Chris Lu c0a9b110dd volume: stop reporting read-only volumes that are no longer here (#10867)
* volume: clear per-collection metrics when a collection leaves a server

The read-only and disk size gauges are only ever set for collections the
heartbeat still finds here, and nothing zeroes the rest. volume.balance marks a
volume read-only to move it, so the last heartbeat that saw it counts it
read-only - and if it was the collection's last volume on that server, that
count stands until the process restarts. The dashboard then shows read-only
volumes that volume.list -readonly cannot find anywhere.

Remember what each heartbeat set, and drop what is gone on the next one.

* volume: stop the read-only volume count from wrapping at 256

The per-collection counters were uint8, so a server holding 256 read-only
volumes of one collection reported zero of them.

* volume: read the read-only flags once when counting them

The heartbeat asked IsReadOnly for the verdict and then read noWriteOrDelete
and noWriteCanDelete straight off the volume, unlocked, so the reasons could
disagree with the verdict they were explaining. Take them together, under one
lock. The location is now nil-checked rather than skipped by short-circuit
evaluation, so a volume that has not joined a disk location yet stays safe.

* volume: let only a surviving volume keep its collection reported

A volume being deleted for expiry still made an entry in the read-only counts,
which is what the cleanup reads as "this collection is still here". The
collection's last volume could go and its series would stand for one more
heartbeat. Count the survivors only.

* volume: size a collection from the volumes it still has

The size totals are rebuilt from scratch every heartbeat, so subtracting a
volume that is about to be deleted took the surviving volumes' sizes down with
it: a collection keeping a small volume and losing a larger one reported the
difference, or lost its entry and kept the previous heartbeat's number.

* volume: cover the deleted bytes total in the surviving volume test

Deleted bytes are totalled the same way as sizes and were going unchecked, so
the test now leaves deleted needles on both volumes and pins that gauge too.
2026-08-21 22:33:01 -07:00
Chris Lu 96304b6870 S3: source config credentials from the environment, and let the chart point at an existing secret (#10868)
* s3: resolve ${VAR} in static config credentials from the environment

A deployment that keeps its S3 keys in a secret store had no way to hand
them to the gateway: -config takes a file, so the keys had to be written
into that file. Let a key in the static config name an environment
variable instead, and drop any credential whose reference stays unset so
the placeholder never becomes a usable key.

* helm: source the generated s3 identities from an existing secret

The only way to reuse credentials that already live in a Secret was to
hand-author the whole seaweedfs_s3_config JSON, since the literal keys in
values.yaml end up in git and a lookup-based keyRef renders empty under
helm template and Argo CD. Let s3.credentials.admin/read name a Secret and
its keys instead: the generated config references them as ${VAR} and the
gateway resolves them from the environment, so nothing is read from the
cluster at render time.

* s3: treat an empty environment value as an unresolved credential reference

A secret store can hand over a key that exists but is blank. Resolving it
would leave an access key whose signing secret is empty, so count it as
unresolved and drop the credential.

* helm: render the s3 secret when only the all-in-one auth flag is set

The all-in-one deployment mounts the s3 secret whenever any of the three
enableAuth flags is set, but the secret itself only rendered for the s3 and
filer flags, so allInOne.s3.enableAuth on its own left the pod waiting on a
secret nothing creates.

* helm ci: check the credential wiring on every workload that mounts it

The render check only looked at the standalone s3 deployment and only at
one of the four variables, so a helper that bound a variable to the wrong
secret key would still pass.

* helm: create the all-in-one s3 secret for every flag that mounts it

The all-in-one pod mounts the secret on any of the three enableAuth flags,
so keying its creation off allInOne.s3.enableAuth alone still left
filer.s3.enableAuth without filer.s3.enabled pointing at a secret nothing
creates. Mirror the deployment's own condition instead, and check each
flag renders both the mount and the secret.

* s3: reject a malformed credential reference instead of keying on it

A typo such as ${MY-VAR} matches no substitution, so it survived expansion
and the placeholder itself became the access key the gateway accepted.
Require every ${ in a static credential to open a well-formed reference.
2026-08-21 22:32:47 -07:00
Chris Lu 35d53a20f6 master: let the leader admit a master that starts with no raft state (#10865)
* master: answer with the leader raft already knows

Topo.Leader() backs off for up to 20 seconds waiting for an election.
Callers that a health probe or a client is blocked on cannot afford that:
/cluster/status, /cluster/healthz and /readyz all sit past the probe
timeout of both the helm chart and the operator, so a master that is
still joining looks dead rather than joining, and the kubelet restarts
it. informNewLeader and SendHeartbeat hold the client on a master that
cannot serve it, exactly when it should move on to find the one that can.

Answer these from MaybeLeader instead, which reports what raft knows
right now. MaybeLeader takes over the "am I the leader myself" fallback
that Leader() used to apply on top of it, so one non-blocking call is
still correct; Leader() keeps the backoff for callers that must wait.

* master: let the leader admit a master that starts with no raft state

Neither raft implementation lets a server outside the configuration
campaign: goraft's promotable() requires a non-empty log, and hashicorp
rejects vote requests from a candidate that is not in its configuration.
A master that comes up with fresh state therefore cannot elect itself in
— the leader has to pull it in. Nothing did.

The peer list is static, rendered from the replica count, so scaling it
up leaves the sitting leader running the old list with no idea the new
masters exist. Under goraft they wait forever. Under hashicorp they are
worse off: each bootstraps a cluster of its own from the new list, and
two of them form a quorum next to the live leader, with their own
TopologyId. That is the split brain SetTopologyId kills a master over.

Admit the peer where it registers instead. Only the leader gets past the
IsLeader check in KeepConnected, and a joining master's client lands
there, so that is the moment it joins. The broadcast OnPeerUpdate rides
on is not enough on its own: it only reaches masters already connected,
which is why a leader that came up first missed both newcomers.

RaftAddServer grew a goraft branch on the way, so cluster.raft.add stops
silently doing nothing on the default raft, and RaftRemoveServer with it.
Bootstrapping is now one call for both implementations, made only after
the peers confirm nobody has a leader, and retried until this master is
in rather than checked once and dropped.

* master: do not evict a peer that is still in -peers

The hashicorp leader drops a master from the raft configuration as soon
as it stops answering pings. A master that is merely restarting answers
nothing, so an ordinary bounce shrinks the quorum behind the operator's
back — and then races its own return: the master comes back, registers,
gets re-admitted, and the eviction lands after it.

A randomized start/stop walk lands on it. Two of three masters running,
the leader evicts the one that just went down, the restart re-adds it,
the removal commits late and takes the leader's own leadership with it.
What is left is a two-server configuration whose other half is down, and
a running master that nobody will ask for a vote — no quorum, no way
back until the third master returns.

-peers is what declares membership. updatePeers already reconciles the
configuration against it on every leadership change, and an operator who
really means to drop a master can say so with cluster.raft.remove, so
keep the eviction for masters that are no longer listed at all.

* test: bounce masters at random and hold the election to it

Twelve rounds of stopping or starting a random master, on both raft
implementations, checking the two things an election must never get
wrong: two masters claiming leadership at once, and a quorum that comes
back without agreeing on one. The cluster's identity has to survive the
whole walk, since a master that re-mints a TopologyId is the split brain
SetTopologyId kills its peers over. The seed is random and logged, so a
failure names the walk that reproduces it.

Below a quorum the walk moves straight on. A master that has lost its
quorum cannot commit anything, and goraft only checks whether it still
has one on an election-timeout ticker, after its peers have been quiet
for a full timeout — measured taking over 30 seconds to step down. That
direction belongs to TestTwoMastersDownAndRestart, which was giving it
ten seconds and would have started failing on a slower machine; it now
waits on that behaviour explicitly rather than sleeping twice and hoping.

WaitForTopologyId returns the id it waited for. Reading it separately
raced the leader applying the raft entry that carries it, which shows up
as an empty id right after an election rather than as a wrong one.
2026-08-21 15:22:22 -07:00
Chris Lu 0c95137528 filer: stop aggregated metadata subscribers from spinning on a peer watermark hold (#10863)
* fix(filer): stop logging a held aggregated read as an error

An aggregated subscriber may not read past the peers' low-watermark, and
it stops at the first entry beyond it by returning a sentinel from the
read callback. LoopProcessLogData logs every callback error, so on a
cluster that keeps writing - where there is almost always an entry newer
than the watermark - every read wrote an ERROR line naming the entry it
stopped at, thousands per minute per filer.

Mark the stop as control flow: an error wrapping StopReadingError is
handed back to the caller unlogged, and the held-read sentinel wraps it.

* fix(filer): release an aggregated watermark hold on peer progress

A held read waited on the aggregated buffer's data channel, which the
next write signalled - but a write cannot release a hold, only a peer
reporting further progress can. On a cluster that keeps writing the loop
therefore re-ran a whole pass per arriving event, log file listing and
all, and held again on the same entry every time.

Signal held readers from the meta aggregator instead, whenever a
low-watermark rises: a peer reporting, or one dropped past its removal
grace. The retry interval stays as the backstop for what no watermark
covers. Count the holds so a parked subscriber stays visible.

* fix(filer): floor how often an aggregated watermark hold releases

Peers advance their delivery watermark on every event they stream, so
releasing a hold on every advance is the same pass-per-event storm as
releasing on every write, just without the log lines - and each pass
lists a day of log files.

Floor the release at 20ms. Advances inside the floor collapse into one
release, which then delivers everything they covered.

* fix(filer): pace a peer's delivery claim by what its subscribers hold at

A filer's local metadata stream carries an idle heartbeat to its peer
aggregators, and each peer turns it into that filer's delivery
low-watermark. Aggregated subscribers hold at the minimum across peers,
so a filer quiet enough to fall back on the heartbeat parked every
subscriber in the cluster up to a keepalive interval - 5 seconds -
behind live writes. With nine filers, most of them quiet at any moment,
the minimum sat there permanently.

Pace that heartbeat at 200ms once the filer has peers. It stays a
keepalive, at the keepalive interval, for a filer with none.

* fix(filer): wake each aggregated hold on its own watermark

A persisted-log read is held by what the peers have flushed, an
in-memory read by what they have delivered, but both parked on one
channel closed whenever either minimum rose. Peers advance their
delivery watermark on every event they stream, so a flush-held reader
woke at the coalescing floor to re-list a day of log files and park
again on the same entry - the storm this set out to fix, in the one
place asymmetric peer progress still reached.

Signal the two separately and park each read on the one that bounds it.
2026-08-21 15:22:05 -07:00
Chris Lu 3bd218e030 volume: cut idle memory at high volume counts (#10861)
* volume: start a volume's batch write worker on first use

Mounting a volume started a goroutine parked on a 128-slot channel, plus
the 128-entry batch slice it had already allocated. That is around 6.7KB
per volume the server pays whether or not the volume ever takes a write:
7231 bytes per mounted volume, of which 4101 is goroutine stack.

Only a write that asks for fsync ever reaches the worker, and a
remote-tiered or read-only volume never can. Create the channel and its
goroutine on the first such request instead, and let a write arriving
after Destroy fall back to the inline path rather than queue onto a
worker that has gone.

Measured over 20000 mounted volumes: 7231 -> 1269 bytes each.

* volume: update the heartbeat report state in place

Every heartbeat built a second map of what it was about to tell the
master, holding a freshly allocated short information message per volume,
then swapped it in over the old one -- and computed departures through a
third map of the live volume ids. A server holding 2M volumes rebuilt all
three every VolumePulsePeriod for a report that usually says nothing.

Number the heartbeats instead and mark the entry already held with the
pass that found the copy, so a quiet volume costs a map lookup and no
allocation. Departures are the entries a pass did not mark; the live-id
map is now built only when there are some, sized to them.

Measured over 10000 mounted volumes: 436 -> 196 bytes allocated per
volume per heartbeat.

* volume: fill one volume information message per heartbeat, not per volume

The heartbeat built a message for every volume held so it could hash it,
then dropped all but the few it had something to say about. At 2M volumes
that is 2M messages allocated every VolumePulsePeriod to send almost none
of them.

Fill a message the caller supplies instead, and replace it only when the
heartbeat keeps it, so a server with nothing to report fills the same one
all the way through.

Measured over 10000 mounted volumes: 196 -> 4 bytes allocated per volume
per heartbeat, and a heartbeat runs a third faster.

* volume: drop the per-volume trace from the heartbeat's status read

glog.V(4).Infof evaluates its arguments whether or not the verbosity is
on, so every volume boxed its id into a fresh interface slice on every
heartbeat: 759 of the 773 allocations a 1000-volume heartbeat made, for a
line that at this scale would print millions of unreadable rows.

Measured over 1000 mounted volumes: 4776 -> 1792 bytes and 759 -> 14
allocations per heartbeat, which no longer grows with the volume count.

* seaweed-volume: mirror the in-place heartbeat report state

Same change as the Go volume server: number the heartbeats and mark the
entry already held with the pass that found the copy, instead of building
a second map of hashes and swapping it in.

The volume snapshot must leave the reporting state as it found it, so it
keeps asking through changed() while a real heartbeat marks through
record().

* volume: refuse writes to a closed volume instead of dereferencing nil

Close and Destroy leave the needle map and data backend nil, but a caller
that already holds the volume can still reach the write path, where both
are used unguarded: a write racing a volume deletion took the server down.
syncDelete has always checked; syncWrite and the batch worker had not.

Reachable before this series and now also from the inline fallback a
durable write takes when the worker has gone.

* seaweed-volume: guard the report state with one mutex, as Go does

The full-list flag and the generation that answers it have to move
together. Split across separate atomics they cannot: a request landing
between begin's two reads returns full == false with the generation it
just raised, and one landing between commit's read and its clear is
marked answered by a heartbeat that carried no list. Either way the
resend is dropped.

Neither is reachable today -- every caller reaches this through the
store's RwLock, the flag setters under a read lock and the heartbeat
build under a write lock, so they cannot interleave. The type should not
depend on that being true two files away, and Go holds a single mutex
over exactly these fields.

* test: build the servers under test to match the harness's offset size

The mixed Go/Rust suites run both servers against one dataset, so both
have to agree on the offset width. They did not: the harness built Go
with no tags, 4-byte offsets, while the Rust crate defaults to its 5bytes
feature, and the Rust server then refused the .vif the Go server had just
written -- "bytes_offset mismatch: found 4, expected 5".

Build each side to match the offset size the test binary itself was
compiled with, so a plain `go test` and one with -tags 5BytesOffset both
get a matched pair.
2026-08-21 13:04:56 -07:00
930603eb74 S3: optionally serve remote-mounted objects from remote when the local read fails (#10837)
* feat(s3): serve from remote on local read failure

When a locally-cached chunk of a remote-mounted object becomes unreadable
(volume server down/restarting, or an evicted needle 404ing under
retry-backoff), fall back to serving the object from its mounted remote
instead of erroring. A bounded pre-flight probe makes a stuck volume trip
the timeout rather than stalling the request.

Gated by -localReadFallbackToRemote (default off) with
-localReadFallbackTimeout (2s default), so existing deployments are
unaffected until they opt in.

* fix(s3): register local-read-fallback flags for mini/server/filer

The mini, server and filer launchers build S3Options directly and only
populate the flag pointers they register. Without registering the two new
flags there, startS3Server dereferenced nil pointers and crashed at boot,
failing every integration suite that runs `weed mini`.

* fix(s3): treat a zero-byte probe read as unreadable

A read that returns no byte -- whether it reports io.EOF or no error at all
-- means the offset is not locally readable, so the probe must fall back to
the remote rather than proceeding to stream a truncated response. Only a
returned byte (including the object's final byte with a trailing io.EOF)
counts as readable.

* s3: finish a mid-stream local read failure from the remote mount

The pre-flight probe only proves the byte at the requested offset readable.
A multi-chunk object can still lose a later chunk after the 200/206 and its
Content-Length are committed, which truncated the body with no fallback.
Resume from the mounted remote at the byte the local copy stopped at, so the
response still carries the declared length. A short local read that surfaces
as a clean EOF is treated the same way instead of silently truncating.

* s3: fall back to the remote mount without a CLI switch

Serving a remote-mounted object from its authoritative remote is what the
read should have done all along -- the alternative is a 500 on an object the
cluster can still reach -- so make it the behavior instead of two new flags,
with the probe bounded by a constant.

* s3: trim the comments on the fallback path

* filer: report only the contiguous prefix when a parallel chunk read fails

The parallel branch of doReadAt fans the chunk reads straight into their own
windows of the output buffer, then sums every task's bytesRead. A middle chunk
failing while a later one succeeds therefore returned a length covering a hole
the reader never filled, handing the caller zeros in the middle of otherwise
valid data.

* s3: only splice the remote onto a local prefix while it is the cached generation

Eligibility establishes a size match, not byte identity: a remote key
overwritten with same-size content between the cache fill and the fallback
would have finished the response with bytes from a second generation, under
the first one's ETag. Stat the remote before resuming and keep the local
error when it no longer matches -- a truncated body is a visible failure,
a spliced one is not.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-21 11:47:53 -07:00