Files
seaweedfs/seaweed-volume/src/remote_storage/mod.rs
T
8ff2e0777e volume server: HTTP DELETE on a distributed EC volume (#11405)
* volume server: HTTP DELETE on a distributed EC volume

The delete handler validated the cookie with EcVolume::read_ec_shard_needle,
which reads only locally-mounted shards and errors "ec shard N not available
locally" for any interval held by a peer. Every Err was mapped to 500 and no
.ecj tombstone was appended, so on a standard 10+4 spread over 14 servers an
HTTP delete of an EC needle could not succeed. The GET path already goes
through read_ec_shard_needle_distributed.

Route the delete's read through the same distributed reader. It does a
local-first pass in its snapshot phase, so the all-shards-local case costs
what it did before, and no store guard is held across the await (the reader
takes its own; RwLockReadGuard is !Send).

Two smaller corrections fall out of the new return type:

  - the reader reports both "needle not in the index" and "volume vanished
    between the has_ec check and the snapshot" as Ok(None), which collapses
    the old Some(Ok(None)) and None arms into one 404;
  - an io::ErrorKind::NotFound now answers 404 rather than 500, matching the
    GET path. Telling a caller to retry a delete that can never succeed was
    half the bug.

The cookie check and its ordering before the journal append are unchanged.

Not addressed here: Rust journals the tombstone locally while Go routes it to
the primary shard holder. That is a separate behaviour change and belongs in
its own PR against the same issue-10 checkbox.

The regression test mounts 13 of 14 shards, leaving out the one holding the
needle's interval. The distributed reader seeds its Reed-Solomon buffers from
locally mounted siblings, so with >= 10 survivors it reconstructs with no peer
fan-out -- which makes the bug reproducible on a single node. Against the
unfixed handler the test fails with 500 vs 202.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* volume server: fail the delete when the EC volume unmounts mid-request

find_ec_volume_mut returning None used to fall through to a 202 with no
.ecj tombstone written, reporting success for a delete that did not
happen. Answer 404 like the other volume-vanished arms so the caller can
retry after a remount.

* volume server: forward EC needle deletes to a primary-shard holder

Mirror Go's doDeleteNeedleFromAtLeastOneRemoteEcShards: the tombstone is
journaled on one holder of the needle's primary data shard via
VolumeEcBlobDelete (or the local journal when this server holds the
shard), falling back to any other shard holder when the primary has
none. Journaling only on the node that received the DELETE scattered
tombstones across whichever server took the request.

* volume server: route BatchDelete EC deletes through the same forwarding

BatchDelete had the same local-journal divergence as HTTP DELETE, plus a
gap the old code admitted in a comment: the .ecx index cannot supply the
needle's cookie, so EC deletes ran with no cookie check at all. A
distributed read now fills the needle for every EC entry — matching Go's
DeleteEcShardNeedle, which reads and compares the fid cookie even when
skip_cookie_check is set — and the tombstone forwards via
delete_ec_shard_needle_distributed. A needle deleted between read and
journal reports 304 like Go's ErrorDeleted; a vanished volume reports
500 so the filer retries.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <devin@cognition.ai>
2026-09-20 23:26:31 -07:00

247 lines
8.6 KiB
Rust

//! Remote storage backends for tiered storage support.
//!
//! Provides a trait-based abstraction over cloud storage providers (S3, GCS, Azure, etc.)
//! and a registry to create clients from protobuf RemoteConf messages.
pub mod endpoint_guard;
pub mod s3;
pub mod s3_tier;
pub use endpoint_guard::{guarded_tcp_connect, validate_remote_endpoint, validate_replica_target};
use crate::pb::remote_pb::{RemoteConf, RemoteStorageLocation};
/// Error type for remote storage operations.
#[derive(Debug, thiserror::Error)]
pub enum RemoteStorageError {
#[error("remote storage type {0} not found")]
TypeNotFound(String),
#[error("remote object not found: {0}")]
ObjectNotFound(String),
#[error("remote storage error: {0}")]
Other(String),
#[error("io error: {0}")]
Io(#[from] std::io::Error),
}
/// Metadata about a remote file entry.
#[derive(Debug, Clone)]
pub struct RemoteEntry {
pub size: i64,
pub last_modified_at: i64, // Unix seconds
pub e_tag: String,
pub storage_name: String,
}
/// Trait for remote storage clients. Matches Go's RemoteStorageClient interface.
#[async_trait::async_trait]
pub trait RemoteStorageClient: Send + Sync {
/// Read (part of) a file from remote storage.
async fn read_file(
&self,
loc: &RemoteStorageLocation,
offset: i64,
size: i64,
) -> Result<Vec<u8>, RemoteStorageError>;
/// Write a file to remote storage.
async fn write_file(
&self,
loc: &RemoteStorageLocation,
data: &[u8],
) -> Result<RemoteEntry, RemoteStorageError>;
/// Get metadata for a file in remote storage.
async fn stat_file(
&self,
loc: &RemoteStorageLocation,
) -> Result<RemoteEntry, RemoteStorageError>;
/// Delete a file from remote storage.
async fn delete_file(&self, loc: &RemoteStorageLocation) -> Result<(), RemoteStorageError>;
/// List all buckets.
async fn list_buckets(&self) -> Result<Vec<String>, RemoteStorageError>;
/// The RemoteConf used to create this client.
fn remote_conf(&self) -> &RemoteConf;
}
/// Create a new remote storage client from a RemoteConf.
pub fn make_remote_storage_client(
conf: &RemoteConf,
) -> Result<Box<dyn RemoteStorageClient>, RemoteStorageError> {
match conf.r#type.as_str() {
// All S3-compatible backends use the same client with different credentials
"s3" | "wasabi" | "backblaze" | "aliyun" | "tencent" | "baidu" | "filebase" | "storj"
| "contabo" => {
let (access_key, secret_key, endpoint, region) = extract_s3_credentials(conf);
Ok(Box::new(s3::S3RemoteStorageClient::new(
conf.clone(),
&access_key,
&secret_key,
&region,
&endpoint,
conf.s3_force_path_style,
)))
}
other => Err(RemoteStorageError::TypeNotFound(other.to_string())),
}
}
/// Endpoint URL that the volume server would dial directly for `conf`, or
/// `None` for non-S3-compatible backends. Every type handled here routes
/// through [`s3::S3RemoteStorageClient`] with a caller-supplied endpoint, so
/// the SSRF guard must validate each one. Keep this match in sync with
/// [`make_remote_storage_client`].
pub fn s3_compatible_endpoint(conf: &RemoteConf) -> Option<&str> {
match conf.r#type.as_str() {
"s3" => Some(&conf.s3_endpoint),
"wasabi" => Some(&conf.wasabi_endpoint),
"backblaze" => Some(&conf.backblaze_endpoint),
"aliyun" => Some(&conf.aliyun_endpoint),
"tencent" => Some(&conf.tencent_endpoint),
"baidu" => Some(&conf.baidu_endpoint),
"filebase" => Some(&conf.filebase_endpoint),
"storj" => Some(&conf.storj_endpoint),
"contabo" => Some(&conf.contabo_endpoint),
_ => None,
}
}
/// Extract S3-compatible credentials from a RemoteConf based on its type.
fn extract_s3_credentials(conf: &RemoteConf) -> (String, String, String, String) {
match conf.r#type.as_str() {
"s3" => (
conf.s3_access_key.clone(),
conf.s3_secret_key.clone(),
conf.s3_endpoint.clone(),
if conf.s3_region.is_empty() {
"us-east-1".to_string()
} else {
conf.s3_region.clone()
},
),
"wasabi" => (
conf.wasabi_access_key.clone(),
conf.wasabi_secret_key.clone(),
conf.wasabi_endpoint.clone(),
conf.wasabi_region.clone(),
),
"backblaze" => (
conf.backblaze_key_id.clone(),
conf.backblaze_application_key.clone(),
conf.backblaze_endpoint.clone(),
conf.backblaze_region.clone(),
),
"aliyun" => (
conf.aliyun_access_key.clone(),
conf.aliyun_secret_key.clone(),
conf.aliyun_endpoint.clone(),
conf.aliyun_region.clone(),
),
"tencent" => (
conf.tencent_secret_id.clone(),
conf.tencent_secret_key.clone(),
conf.tencent_endpoint.clone(),
String::new(),
),
"baidu" => (
conf.baidu_access_key.clone(),
conf.baidu_secret_key.clone(),
conf.baidu_endpoint.clone(),
conf.baidu_region.clone(),
),
"filebase" => (
conf.filebase_access_key.clone(),
conf.filebase_secret_key.clone(),
conf.filebase_endpoint.clone(),
String::new(),
),
"storj" => (
conf.storj_access_key.clone(),
conf.storj_secret_key.clone(),
conf.storj_endpoint.clone(),
String::new(),
),
"contabo" => (
conf.contabo_access_key.clone(),
conf.contabo_secret_key.clone(),
conf.contabo_endpoint.clone(),
conf.contabo_region.clone(),
),
_ => (
conf.s3_access_key.clone(),
conf.s3_secret_key.clone(),
conf.s3_endpoint.clone(),
conf.s3_region.clone(),
),
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn s3_compatible_endpoint_covers_all_s3_backends() {
let s3 = RemoteConf {
r#type: "s3".to_string(),
s3_endpoint: "http://s3.internal".to_string(),
..Default::default()
};
assert_eq!(s3_compatible_endpoint(&s3), Some("http://s3.internal"));
// A non-"s3" S3-compatible type still surfaces its own endpoint, so the
// SSRF guard cannot be bypassed by picking a different alias.
let wasabi = RemoteConf {
r#type: "wasabi".to_string(),
wasabi_endpoint: "http://wasabi.internal".to_string(),
..Default::default()
};
assert_eq!(
s3_compatible_endpoint(&wasabi),
Some("http://wasabi.internal")
);
// Non-S3 backends do not dial a caller-supplied URL directly.
let gcs = RemoteConf {
r#type: "gcs".to_string(),
..Default::default()
};
assert_eq!(s3_compatible_endpoint(&gcs), None);
}
#[test]
fn azure_endpoint_has_no_ssrf_path() {
// The Go volume server guards the caller-supplied azure endpoint against
// SSRF. This server has no azure backend, so there is nothing to dial:
// azure is not S3-compatible (the endpoint guard does not apply) and
// make_remote_storage_client rejects the type before building a client.
let azure = RemoteConf {
r#type: "azure".to_string(),
azure_endpoint: "https://169.254.169.254/".to_string(),
..Default::default()
};
assert_eq!(s3_compatible_endpoint(&azure), None);
assert!(make_remote_storage_client(&azure).is_err());
}
#[test]
fn gcs_credentials_have_no_ssrf_path() {
// The Go volume server accepts only static-key gcs credentials and puts
// their token endpoint behind the SSRF guard, because the SDK dials
// whatever url, file or executable the credentials name. This server has
// no gcs backend, so make_remote_storage_client rejects the type before
// any credentials are parsed. Anyone adding one must carry both guards
// over with it.
let gcs = RemoteConf {
r#type: "gcs".to_string(),
gcs_google_application_credentials: r#"{"type":"external_account","credential_source":{"url":"http://169.254.169.254/"}}"#.to_string(),
..Default::default()
};
assert_eq!(s3_compatible_endpoint(&gcs), None);
assert!(make_remote_storage_client(&gcs).is_err());
}
}