Files
seaweedfs/seaweed-volume/src/server/volume_server.rs
T
07da302da0 volume server: ec.decode verifies, cleans up and compacts like Go, off the runtime (#11547)
* volume server: ec.decode reads the .ecx from the index dir it was copied to

VolumeEcShardsCopy writes the .ecx/.ecj into the receiver's -dir.idx, so
with a split data/index dir the decode target has no .ecx beside its
shards. VolumeEcShardsToVolume sized the .dat from the right .ecx but
built the .idx from the data dir, failing with NotFound after the .dat
was already published. It now reads .ecx/.ecj from where the EC volume
opened them and writes the .idx beside the .dat, where Go leaves it.

The live-entry check and the .dat size also ignored deletions recorded
only in the .ecj, which Go folds into the .ecx (RebuildEcxFile) first:
a fully deleted volume was decoded instead of reported as having no live
entries, and deleted tail needles were copied into the .dat. Both now
treat journaled ids as deleted, without rewriting the sealed .ecx.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps the decoded volume writable and reads every .ecj

The rebuilt .idx copied a journaled tail needle's .ecx row verbatim after
the .dat was cut short before it, so the mount saw a row past EOF and
marked the decoded volume read-only. Rows of deleted needles the .dat no
longer holds are now dropped, and each journaled needle still in the .dat
gets one tombstone instead of one per journal entry.

VolumeEcShardsCopy appends journals collected from other holders into
the idx dir, but the decode read only the .ecj beside the .ecx, which
sits in the data dir when this server generated the shards. It now
reads both, once, in bounded chunks via the loader EcVolume uses.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: test ec.decode drops a sealed .ecx tail tombstone

Covers the other half of the rule added in the previous commit: a tail
needle tombstoned in the .ecx itself (Go's RebuildEcxFile) is cut from
the .dat, and its row must not reach the rebuilt .idx either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode runs its file I/O off the async runtime

VolumeEcShardsToVolume released the store lock before decoding, but read
the .ecx/.ecj, rebuilt the .dat and wrote the .idx inside the async
handler, parking a runtime worker for the length of a volume-sized copy.
The decode now runs in spawn_blocking on inputs snapshotted under the
store lock.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode checks the rebuilt .dat is complete

Go stats the decoded .dat before writing the .idx (VerifyDecodedDatFile)
and fails the decode when it is shorter than the extent the EC index
references, since the caller deletes the shards once the call returns.
The Rust handler returned success without that check. The rebuild
already fails on a short shard read, so this guards the published file
itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode drops the decoded volume's bitrot sidecars

Go removes <base>.ecsum and <base>.ecsum.v<N> beside the .dat and beside
the .ecx once the .idx is written, so a stale checksum sidecar cannot
pass for the protection of a later re-encode. The Rust handler left them
in place. Removal is best effort, as in Go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode compacts the decoded volume

Go ends VolumeEcShardsToVolume with an offline CompactVolumeFiles, so the
decoded volume holds only live needles. The Rust decode left every needle
deleted through the .ecj in the .dat, tombstoned in the .idx, until a
later vacuum reclaimed it.

Store::compact_volume_files loads the unmounted volume, checks free space
the way the vacuum does (the estimate now lives in one helper), and runs
the vacuum's compact-by-index and commit. As in Go a failed compaction is
logged and the decode still succeeds, so the uncompacted .idx rules stay:
the tests that pin them now make the compaction fail.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps deletes journaled while the .dat is written

The decode read the .ecj journals once, before rebuilding the .dat, so a
delete that reached the EC volume during the rebuild was left out of the
new .idx and the needle came back live. Each journal's read length is now
kept, and the bytes appended since are read just before the .idx is
written, after waiting out any journal append in flight (appends hold
the store write lock), so every delete acknowledged by then is in the
.idx. A delete after that point is still lost, as in Go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Guard overlapping ec decode requests; serialize journal catch-up

volume_ec_shards_to_volume runs its decode in spawn_blocking, so a
dropped request leaves the job running and a retry would race it on the
temporary and final volume files. Claim the vid in a per-server
in-flight set until the blocking job finishes, and return Unavailable
to an overlapping request. The Go handler has the same exposure and
gets the same guard.

Journal appends hold the store write lock through their
sync-or-truncate, so holding a read lock across the catch-up read
guarantees every record it sees is committed: a rolled-back delete can
no longer leave a tombstone in the decoded index.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Reconcile the swap when offline compaction commit fails

A CommitCompact that fails after the .cpc marker may have renamed .dat
but not .idx. cleanup_compact refuses while the marker exists, so the
mismatched pair survived until a restart reconciled it — and the decode
caller treats the failure as non-fatal. Run reconcileCompactState on
commit failure so a decided swap rolls forward and orphan temps are
removed before the volume can mount.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Release the decode claim on panic

* volume: add ec_decodes_in_flight to the integration-test state literal

* volume server: hold the decode tail's lock through compaction

The catch_up read released before the rebuilt .idx was written and the
volume compacted, so a delete synced to .ecj in that window was durably
journaled yet absent from the published index — resurrecting the needle.
Rust now holds the store read lock from catch_up through compact, and Go
mirrors it by holding the volume's journal lock from the journal-
consuming index write through CompactVolumeFiles.

* volume server: serialize ec decode's tail per volume, not per store

Review follow-ups on the decode path:

- Rust: holding the store read lock from journal catch-up through the
  offline compaction stalled every writer on unrelated volumes for the
  whole rewrite. The new ec_decode_tail set marks the vid only while its
  .idx is published and .cpd/.cpx swapped; the two local .ecj append paths
  (VolumeEcBlobDelete, the distributed delete's local journal) wait on a
  Notify for that span — Go's per-volume ecjFileAccessLock semantics
  without the global stall. VolumeMount and the staged-adopt path are also
  held off while a decode claim is in flight so neither can race the swap.

- Rust: the initial journal read ran unlocked, so bytes a rolled-back
  append later truncated could be folded in as phantom tombstones. The
  first pass stays unlocked (a slow journal must not stall the store) and
  a rescan under the quiescing read lock re-reads only committed content;
  catch_up now rebuilds the id set when a regular journal shrank.

- Go: the decode resolved the compaction DiskLocation through
  FindEcVolume while holding the journal lock, inverting DestroyEcVolume's
  map->journal order into a deadlock. The lookup now happens first, and
  DestroyEcVolume/deleteEcVolumeById/DiskLocation.Close destroy outside
  the map lock.

- Go: RebuildEcxFile unlinks .ecj while the volume's ecjFile handle stays
  open, so later deletes could commit to a detached inode. Both call sites
  now fold under the journal lock and ReopenDeletionJournal repoints the
  handle at the live path, working on the volume's resolved .ecx dir
  (EcIndexBaseFileName) rather than the configured index dir.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: fence EC remounts behind the destroy tombstone

DestroyEcVolume, deleteEcVolumeById, and the collection-delete sweep now
remove the EcVolume from ecVolumes before destroying it off-lock, so a
concurrent remount could re-open shard files that the in-flight destroy
then unlinks — registering a detached fd.

Each destroy records a per-vid tombstone channel in a new
ecVolumesDestroying map before dropping the map entry and closes it when
Destroy returns. The tombstone intentionally survives as the vid's
destroy generation: loadEcShardWithIdxDir compares it before and after
opening the shard, so a destroy that both started and finished inside the
open window is still detected. A mismatch drops the just-opened shard
(releasing its fd and mount gauge) and retries after the destroy
completes; a successful mount clears the stale tombstone.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: rescan the .ecj under the store lock only after a rollback

The decode's second journal pass ran a full rescan under the store read
lock on every decode, stalling unrelated writers for the length of the
scan. Bump a process-wide epoch whenever a failed append truncates its
uncommitted tail; an unchanged epoch between the unlocked read and the
quiesced pass proves every id folded in was committed, so catch_up()
suffices. catch_up() also treats a journal that was read but has since
disappeared as shrunk to zero, so its earlier ids cannot linger.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: check the decode tail under the store write lock on delete

A blob delete waited for the publishing tail before taking the store
write lock, so a decode that claimed the tail while the delete was
parked behind the decoder's read lock could still see the journal append
land after the rebuilt .idx — an acknowledged delete the mount would
miss. Test tail membership under the write lock instead, retrying after
the wait; journal_delete_local reports WouldBlock for the same recheck
on the distributed path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: claim the vid for mount and staged adoption, per volume

VolumeMount and the staged .copying adoption held the
ec_decodes_in_flight set lock through slow file renames and mounts,
stalling every unrelated volume's decode, mount, and adoption. Take the
per-volume claim instead — the same exclusion against a racing decode
for this vid, released when the call returns.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: fail the decode when a compaction commit marker survives

CompactVolumeFiles' caller logged a compaction error and went on to
delete the EC shards. When the commit marker (.cpc) is still on disk the
.dat/.idx swap was decided but could not be reconciled, so the mounted
pair may be mismatched — report the failure instead so the shards are
kept and the caller can retry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: gate the parked-delete test on the held write lock

The releaser thread and the spawned delete raced for the store write
lock; on a slow runner the delete could acquire it first and commit
before the tail was ever claimed, failing !delete.is_finished() on the
Windows unit-test job. Spawn the delete only after the thread reports
the lock held.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:55:15 +08:00

556 lines
22 KiB
Rust

//! VolumeServer: the main HTTP server for volume operations.
//!
//! Routes:
//! GET/HEAD /{vid},{fid} — read a file
//! POST/PUT /{vid},{fid} — write a file
//! DELETE /{vid},{fid} — delete a file
//! GET /status — server status
//! GET /healthz — health check
//!
//! Matches Go's server/volume_server.go.
use std::net::SocketAddr;
use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32, Ordering};
use std::sync::{Arc, RwLock};
use axum::{
Router,
extract::{Request, State, connect_info::ConnectInfo},
http::{HeaderValue, Method, StatusCode, header},
middleware::{self, Next},
response::{IntoResponse, Response},
routing::{any, get},
};
use crate::config::ReadMode;
use crate::security::Guard;
use crate::storage::store::Store;
use super::grpc_client::OutgoingGrpcTlsConfig;
use super::handlers;
use super::write_queue::WriteQueue;
#[derive(Clone, Debug, Default)]
pub struct RuntimeMetricsConfig {
pub push_gateway: crate::metrics::PushGatewayConfig,
}
/// Shared state for the volume server.
pub struct VolumeServerState {
pub store: RwLock<Store>,
pub guard: RwLock<Guard>,
pub is_stopping: RwLock<bool>,
/// Maintenance mode flag.
pub maintenance: AtomicBool,
/// State version — incremented on each SetState call.
pub state_version: AtomicU32,
/// Throttling: concurrent upload/download limits (in bytes, 0 = disabled).
pub concurrent_upload_limit: i64,
pub concurrent_download_limit: i64,
pub inflight_upload_data_timeout: std::time::Duration,
pub inflight_download_data_timeout: std::time::Duration,
/// Current in-flight upload/download bytes.
pub inflight_upload_bytes: AtomicI64,
pub inflight_download_bytes: AtomicI64,
/// Notify waiters when inflight bytes decrease.
pub upload_notify: tokio::sync::Notify,
pub download_notify: tokio::sync::Notify,
/// Data center name from config.
pub data_center: String,
/// Rack name from config.
pub rack: String,
/// File size limit in bytes (0 = no limit).
pub file_size_limit_bytes: i64,
/// Default IO rate limit for maintenance copy/replication work.
pub maintenance_byte_per_second: i64,
/// Whether the server is connected to master (heartbeat active).
pub is_heartbeating: AtomicBool,
/// Whether master addresses are configured.
pub has_master: bool,
/// Seconds to wait before shutting down servers (graceful drain).
pub pre_stop_seconds: u32,
/// Notify heartbeat to send an immediate update when volume state changes.
pub volume_state_notify: tokio::sync::Notify,
/// Optional batched write queue for improved throughput under load.
pub write_queue: std::sync::OnceLock<WriteQueue>,
/// Read mode: local, proxy, or redirect for non-local volumes.
pub read_mode: ReadMode,
/// If true, FetchAndWriteNeedle skips remote S3 endpoint validation,
/// allowing arbitrary (incl. loopback / link-local / metadata) hosts.
pub allow_untrusted_remote_endpoints: bool,
/// First master address for volume lookups (e.g., "localhost:9333").
pub master_url: String,
/// Seed master addresses for UI rendering.
pub master_urls: Vec<String>,
/// Canonical http `host:port` form of every configured seed master.
/// Built once at construction so Ping admission stays O(1). Mirrors
/// Go's `seedMasterSet` on `VolumeServer`.
pub seed_master_set: std::collections::HashSet<String>,
/// Current master this server is heartbeating with, in canonical http
/// `host:port` form. Empty when no heartbeat connection is active. The
/// heartbeat goroutine writes; admission reads — the lock keeps them
/// from racing on a leader change. Mirrors Go's `currentMaster` plus
/// `currentMasterLock`.
pub current_master_url: tokio::sync::RwLock<String>,
/// This server's own address (ip:port) for filtering self from lookup results.
pub self_url: String,
/// HTTP client for proxy requests and master lookups.
pub http_client: reqwest::Client,
/// Scheme used for outgoing master and peer HTTP requests ("http" or "https").
pub outgoing_http_scheme: String,
/// Optional client TLS material for outgoing gRPC connections.
pub outgoing_grpc_tls: Option<OutgoingGrpcTlsConfig>,
/// Metrics push settings learned from master heartbeat responses.
pub metrics_runtime: std::sync::RwLock<RuntimeMetricsConfig>,
pub metrics_notify: tokio::sync::Notify,
/// Whether JPEG uploads should be normalized using EXIF orientation.
pub fix_jpg_orientation: bool,
/// Read tuning flags for large-file streaming.
pub has_slow_read: bool,
pub read_buffer_size_bytes: usize,
/// Path to security.toml — stored for SIGHUP reload.
pub security_file: String,
/// Original CLI whitelist entries — stored for SIGHUP reload.
pub cli_white_list: Vec<String>,
/// Path to state.pb file for persisting VolumeServerState across restarts.
pub state_file_path: String,
/// Volumes with an EC decode in flight. A dropped request leaves the
/// blocking job running; this keeps a retry from racing it on the
/// same volume files.
pub ec_decodes_in_flight:
std::sync::Mutex<std::collections::HashSet<crate::storage::types::VolumeId>>,
/// Volumes whose EC decode is in its publishing tail (journal catch-up,
/// .idx write, compaction). Local .ecj appenders wait on
/// `ec_decode_tail_notify` while their vid is listed, so no committed
/// delete falls between the last catch_up and the .cpd/.cpx swap —
/// the per-volume slice of Go's EcVolume.ecjFileAccessLock.
pub ec_decode_tail:
std::sync::Mutex<std::collections::HashSet<crate::storage::types::VolumeId>>,
/// Wakes .ecj appenders waiting on `ec_decode_tail` when a decode's
/// publishing tail ends.
pub ec_decode_tail_notify: tokio::sync::Notify,
}
impl VolumeServerState {
/// Check if the server is in maintenance mode; return gRPC error if so.
pub fn check_maintenance(&self) -> Result<(), tonic::Status> {
if self.maintenance.load(Ordering::Relaxed) {
let id = self.store.read().unwrap().id.clone();
return Err(tonic::Status::unavailable(format!(
"volume server {} is in maintenance mode",
id
)));
}
Ok(())
}
/// Build the seed master set from a list of raw `host:port[.grpcPort]`
/// addresses, normalised the same way Go's `pb.ServerAddress.ToHttpAddress`
/// does (drop the `.grpcPort` suffix, preserve everything else).
pub fn build_seed_master_set(master_urls: &[String]) -> std::collections::HashSet<String> {
master_urls
.iter()
.map(|m| to_http_address(m).into_owned())
.collect()
}
/// Returns true iff `target` (normalised to canonical http `host:port`)
/// is a master this server already knows about. Volume servers do not
/// keep a peer-volume or peer-filer list, so Ping is scoped to masters.
/// Mirrors Go's `VolumeServer.isKnownPingTarget`.
pub async fn is_known_ping_target(&self, target: &str, target_type: &str) -> bool {
if target_type != "master" {
return false;
}
let key = to_http_address(target).into_owned();
if key.is_empty() {
return false;
}
let current = self.current_master_url.read().await.clone();
if !current.is_empty() && current == key {
return true;
}
self.seed_master_set.contains(&key)
}
}
pub fn build_metrics_router() -> Router {
Router::new().route("/metrics", get(handlers::metrics_handler))
}
pub fn normalize_outgoing_http_url(scheme: &str, raw_target: &str) -> Result<String, String> {
if raw_target.starts_with("http://") || raw_target.starts_with("https://") {
let mut url = reqwest::Url::parse(raw_target)
.map_err(|e| format!("invalid url {}: {}", raw_target, e))?;
url.set_scheme(scheme)
.map_err(|_| format!("invalid scheme {}", scheme))?;
return Ok(url.to_string());
}
Ok(format!("{}://{}", scheme, raw_target))
}
/// Convert a SeaweedFS server address to its HTTP `host:port` form.
///
/// Mirrors Go's `pb.ServerAddress.ToHttpAddress()`: SeaweedFS encodes a
/// server's gRPC port by appending `.grpcPort` to the HTTP port (e.g.
/// `host:9333.19333`). For HTTP requests we want only the HTTP `host:port`.
/// - `host:port.grpcPort` -> `host:port`
/// - `host:port` -> `host:port` (unchanged)
/// - Anything that does not look like `host:port[.grpcPort]` is returned unchanged.
///
/// Returns a `Cow<str>` so the common (no-suffix) case borrows from `addr`
/// without allocating; only the rewrite branch produces a new `String`.
pub fn to_http_address(addr: &str) -> std::borrow::Cow<'_, str> {
let Some(ports_sep_index) = addr.rfind(':') else {
return std::borrow::Cow::Borrowed(addr);
};
let ports = &addr[ports_sep_index + 1..];
if let Some(dot_idx) = ports.rfind('.') {
let http_port = &ports[..dot_idx];
let grpc_port = &ports[dot_idx + 1..];
// Only strip the suffix when both parts parse as real ports — leave
// anything else (e.g. "host:abc.def") untouched so bad config surfaces
// rather than being silently rewritten. Mirrors the validation already
// done in `to_grpc_address` for the inverse direction.
if let (Ok(_), Ok(_)) = (http_port.parse::<u16>(), grpc_port.parse::<u16>()) {
return std::borrow::Cow::Owned(addr[..ports_sep_index + 1 + dot_idx].to_string());
}
}
std::borrow::Cow::Borrowed(addr)
}
fn request_remote_addr(request: &Request) -> Option<SocketAddr> {
request
.extensions()
.get::<ConnectInfo<SocketAddr>>()
.map(|info| info.0)
}
fn request_is_whitelisted(state: &VolumeServerState, request: &Request) -> bool {
request_remote_addr(request)
.map(|remote_addr| {
state
.guard
.read()
.unwrap()
.check_whitelist(&remote_addr.to_string())
})
.unwrap_or(true)
}
/// Middleware: set Server header, echo x-amz-request-id, set CORS if Origin present.
async fn common_headers_middleware(request: Request, next: Next) -> Response {
let origin = request.headers().get("origin").cloned();
let request_id = super::request_id::generate_http_request_id();
let mut response =
super::request_id::scope_request_id(
request_id.clone(),
async move { next.run(request).await },
)
.await;
let headers = response.headers_mut();
if let Ok(val) = HeaderValue::from_str(crate::version::server_header()) {
headers.insert("Server", val);
}
if let Ok(val) = HeaderValue::from_str(&request_id) {
headers.insert("X-Request-Id", val.clone());
headers.insert("x-amz-request-id", val);
}
if origin.is_some() {
headers.insert("Access-Control-Allow-Origin", HeaderValue::from_static("*"));
headers.insert(
"Access-Control-Allow-Credentials",
HeaderValue::from_static("true"),
);
}
response
}
/// Admin store handler — dispatches based on HTTP method.
/// Matches Go's privateStoreHandler: GET/HEAD → read, POST/PUT → write,
/// DELETE → delete, OPTIONS → CORS headers, anything else → 400.
async fn admin_store_handler(state: State<Arc<VolumeServerState>>, request: Request) -> Response {
let start = std::time::Instant::now();
let method = request.method().clone();
let mut method_str = method.as_str().to_string();
let request_bytes = request
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
.unwrap_or(0);
super::server_stats::record_request_open();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
.with_label_values(&[&method_str])
.inc();
let whitelist_rejected = matches!(method, Method::POST | Method::PUT | Method::DELETE)
&& !request_is_whitelisted(&state, &request);
let response = match method.clone() {
_ if whitelist_rejected => StatusCode::UNAUTHORIZED.into_response(),
Method::GET | Method::HEAD => {
super::server_stats::record_read_request();
handlers::get_or_head_handler_from_request(state, request).await
}
Method::POST | Method::PUT => {
super::server_stats::record_write_request();
if request_bytes > 0 {
super::server_stats::record_bytes_in(request_bytes);
}
handlers::post_handler(state, request).await
}
Method::DELETE => {
super::server_stats::record_delete_request();
handlers::delete_handler(state, request).await
}
Method::OPTIONS => {
super::server_stats::record_read_request();
admin_options_response()
}
_ => {
let method_name = request.method().to_string();
let query = request.uri().query().map(|q| q.to_string());
method_str = "INVALID".to_string();
handlers::json_error_with_query(
StatusCode::BAD_REQUEST,
format!("unsupported method {}", method_name),
query.as_deref(),
)
}
};
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
.with_label_values(&[&method_str])
.dec();
crate::metrics::REQUEST_COUNTER
.with_label_values(&[&method_str, response.status().as_str()])
.inc();
crate::metrics::REQUEST_DURATION
.with_label_values(&[&method_str])
.observe(start.elapsed().as_secs_f64());
response
}
/// Public store handler — dispatches based on HTTP method.
/// Matches Go's publicReadOnlyHandler: GET/HEAD → read, OPTIONS → CORS,
/// anything else → 200 (passthrough no-op).
async fn public_store_handler(state: State<Arc<VolumeServerState>>, request: Request) -> Response {
let start = std::time::Instant::now();
let method = request.method().clone();
let method_str = method.as_str().to_string();
super::server_stats::record_request_open();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
.with_label_values(&[&method_str])
.inc();
let response = match method.clone() {
Method::GET | Method::HEAD => {
super::server_stats::record_read_request();
handlers::get_or_head_handler_from_request(state, request).await
}
Method::OPTIONS => {
super::server_stats::record_read_request();
public_options_response()
}
_ => StatusCode::OK.into_response(),
};
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
.with_label_values(&[&method_str])
.dec();
crate::metrics::REQUEST_COUNTER
.with_label_values(&[&method_str, response.status().as_str()])
.inc();
crate::metrics::REQUEST_DURATION
.with_label_values(&[&method_str])
.observe(start.elapsed().as_secs_f64());
response
}
/// Build OPTIONS response for admin port.
fn admin_options_response() -> Response {
let mut response = StatusCode::OK.into_response();
let headers = response.headers_mut();
headers.insert(
"Access-Control-Allow-Methods",
HeaderValue::from_static("PUT, POST, GET, DELETE, OPTIONS"),
);
headers.insert(
"Access-Control-Allow-Headers",
HeaderValue::from_static("*"),
);
response
}
/// Build OPTIONS response for public port.
fn public_options_response() -> Response {
let mut response = StatusCode::OK.into_response();
let headers = response.headers_mut();
headers.insert(
"Access-Control-Allow-Methods",
HeaderValue::from_static("GET, OPTIONS"),
);
headers.insert(
"Access-Control-Allow-Headers",
HeaderValue::from_static("*"),
);
response
}
/// Build the admin (private) HTTP router — supports all operations.
/// UI route is only registered when no signing keys are configured,
/// matching Go's `if signingKey == "" || enableUiAccess` check.
pub fn build_admin_router(state: Arc<VolumeServerState>) -> Router {
let guard = state.guard.read().unwrap();
// This helper can only derive the default Go behavior from the guard state:
// UI stays enabled when the write signing key is empty. The explicit
// `access.ui` override is handled by `build_admin_router_with_ui(...)`.
let ui_enabled = guard.signing_key.0.is_empty();
drop(guard);
build_admin_router_with_ui(state, ui_enabled)
}
/// Build the admin router with an explicit UI exposure flag.
pub fn build_admin_router_with_ui(state: Arc<VolumeServerState>, ui_enabled: bool) -> Router {
let mut router = Router::new()
.route("/status", get(handlers::status_handler))
.route("/healthz", get(handlers::healthz_handler))
.route("/favicon.ico", get(handlers::favicon_handler))
.route(
"/seaweedfsstatic/{*path}",
get(handlers::static_asset_handler),
)
.route("/", any(admin_store_handler))
.route("/{path}", any(admin_store_handler))
.route("/{vid}/{fid}", any(admin_store_handler))
.route("/{vid}/{fid}/{filename}", any(admin_store_handler))
.fallback(admin_store_handler);
if ui_enabled {
// Note: /stats/* endpoints are commented out in Go's volume_server.go (L130-134).
// Only the UI endpoint is registered when UI access is enabled.
router = router.route("/ui/index.html", get(handlers::ui_handler));
}
router
.layer(middleware::from_fn(common_headers_middleware))
.with_state(state)
}
/// Build the public (read-only) HTTP router — only GET/HEAD.
pub fn build_public_router(state: Arc<VolumeServerState>) -> Router {
Router::new()
.route("/favicon.ico", get(handlers::favicon_handler))
.route(
"/seaweedfsstatic/{*path}",
get(handlers::static_asset_handler),
)
.route("/", any(public_store_handler))
.route("/{path}", any(public_store_handler))
.route("/{vid}/{fid}", any(public_store_handler))
.route("/{vid}/{fid}/{filename}", any(public_store_handler))
.fallback(public_store_handler)
.layer(middleware::from_fn(common_headers_middleware))
.with_state(state)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_to_http_address_strips_grpc_port_suffix() {
assert_eq!(to_http_address("10.0.0.1:9333.19333"), "10.0.0.1:9333");
assert_eq!(
to_http_address("master.local:9333.19333"),
"master.local:9333"
);
assert_eq!(to_http_address("10.85.183.6:5300.6300"), "10.85.183.6:5300");
}
#[test]
fn test_to_http_address_passthrough_without_grpc_suffix() {
assert_eq!(to_http_address("10.0.0.1:9333"), "10.0.0.1:9333");
assert_eq!(to_http_address("master.local:9333"), "master.local:9333");
}
#[test]
fn test_to_http_address_returns_input_when_unparseable() {
assert_eq!(to_http_address(""), "");
assert_eq!(to_http_address("no-port"), "no-port");
// Trailing colon: nothing after the separator, treat as unparseable.
assert_eq!(to_http_address("host:"), "host:");
}
#[test]
fn test_to_http_address_borrows_when_unchanged_and_owns_when_stripped() {
// The common case (no suffix) must not allocate.
let result = to_http_address("10.0.0.1:9333");
assert!(matches!(result, std::borrow::Cow::Borrowed(_)));
// Stripping requires a new string.
let result = to_http_address("10.0.0.1:9333.19333");
assert!(matches!(result, std::borrow::Cow::Owned(_)));
assert_eq!(result, "10.0.0.1:9333");
// Unparseable / passthrough also borrows.
let result = to_http_address("host:abc.def");
assert!(matches!(result, std::borrow::Cow::Borrowed(_)));
}
#[test]
fn test_to_http_address_keeps_non_numeric_dotted_suffix() {
// The dotted form is only valid when both sides are real port numbers.
// Otherwise the address is malformed config (e.g. a hostname like
// "host:abc.def"), and silently rewriting it would just hide the bug.
assert_eq!(to_http_address("host:abc.def"), "host:abc.def");
assert_eq!(to_http_address("host:9333.notaport"), "host:9333.notaport");
assert_eq!(
to_http_address("host:notaport.19333"),
"host:notaport.19333"
);
// Out-of-range ports must not be silently truncated either.
assert_eq!(to_http_address("host:99999.19333"), "host:99999.19333");
}
#[test]
fn test_to_http_address_handles_bracketed_ipv6_literals() {
// The function uses `rfind(':')`, so for bracketed IPv6 the port
// separator is correctly identified as the colon AFTER the closing
// bracket — making IPv4 and IPv6 behave the same.
assert_eq!(to_http_address("[::1]:9333.19333"), "[::1]:9333");
assert_eq!(
to_http_address("[2001:db8::10]:5300.6300"),
"[2001:db8::10]:5300"
);
// Plain bracketed IPv6 without a dotted suffix is borrowed unchanged.
let result = to_http_address("[2001:db8::1]:9333");
assert!(matches!(result, std::borrow::Cow::Borrowed(_)));
assert_eq!(result, "[2001:db8::1]:9333");
// Non-numeric / out-of-range / missing suffix all preserve the input.
assert_eq!(to_http_address("[::1]:"), "[::1]:");
assert_eq!(to_http_address("[::1]:abc.def"), "[::1]:abc.def");
assert_eq!(to_http_address("[::1]:99999.19333"), "[::1]:99999.19333");
}
}