rust: clippy-clean both crates and adopt the std APIs the 1.91 MSRV allows (#11312)

* rust: apply clippy --fix to both crates

The mechanical part of a clippy sweep: `cargo clippy --all-targets --fix`
on seaweed-volume and the seaweed-worker workspace, hand-reviewed. Both
manifests declare their MSRV (1.91.1 and 1.94.1), so every suggestion
clippy applied is within it: the collapsible_if sites become let chains
(1.88, edition 2024), `% n == 0` becomes is_multiple_of (1.87),
chunks_exact with a constant becomes as_chunks (1.88), repeat().take()
becomes repeat_n (1.82), and io::Error::new(Other, ..) becomes
io::Error::other (1.74). The rest is redundant clones, borrows, casts,
closures and field names.

Nothing here changes behaviour. The three let_and_return sites in
needle_map.rs and store_ec.rs deserve a note: the `let result = ..;
result` shape was a deliberate edition-2021 workaround to drop a redb
guard before the table it borrows. Edition 2024 drops tail-expression
temporaries before locals, which is why clippy now flags it, and the
two comments that described the workaround say so instead.

Manual edits on top of the tool output: the blocks clippy rewrote are
re-indented the way rustfmt lays them out (only those blocks — the
crate is not rustfmt-clean and a whole-crate fmt would bury this diff),
the blank lines let_and_return left behind are removed, and the CRC
legacy_value test compares against a literal worked out from the
original shift formula rather than restating rotate_right.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: clear the clippy warnings --fix cannot apply, and say why the rest stay

Hand fixes for the lints clippy only reports. Behaviour is unchanged
throughout; each rewrite is the one clippy names.

- needless_range_loop (7): index loops over shard vectors become
  iterator loops. Where the old code indexed `v[..n]` the new loop
  iterates `v[..n]` so an undersized vector still panics the same way.
- field_reassign_with_default (6): struct literals with `..Default`.
- redundant_pattern_matching (3): `if let Err(_) = guard.check()` becomes
  `.is_err()`, which also releases the read guard at the end of the
  condition instead of at the end of the block.
- manual_strip (2), manual_checked_ops, format_in_format_args,
  redundant_locals, wrong_self_convention (to_vif takes self by value,
  so it is into_vif; CompactEntry is Copy, so to_needle_value takes self).
- type_complexity (2): `OrphanShardLoad` and `RawNeedleEntry` name two
  tuples that were spelled out inline.
- new_without_default: CompactNeedleMap gets a Default that calls new().
- suspicious_open_options: a test helper spells out `.truncate(false)`,
  which is what `.create(true).write(true)` already did.

What stays, and the attribute that says so:

- too_many_arguments (10): `#[expect]` on each function. Folding 8–15
  parameters into a struct is a design change, not a lint fix.
- await_holding_lock / readonly_write_lock: one test holds the store
  write guard across a sleep on purpose, as a barrier that parks the
  copy task at the mount block. `#[expect(.., reason = ..)]` records it.
- module_inception: needle/needle.rs mirrors the Go package layout.

Two lints become crate-wide policy in `[lints.clippy]`, with the reason
next to each: result_large_err, because every RPC path returns
tonic::Status (176 bytes) and boxing it would change every handler
signature; and needless_update, because `..Default::default()` on a
protobuf message literal is what lets a proto gain a field without
touching every constructor (all 11 sites are pb messages). The worker
workspace gets the same table and its members opt in with
`lints.workspace = true`; its generated plugin.rs also allows
large_enum_variant on prost's oneof enums.

Both crates are now clean under `cargo clippy --all-targets -- -D warnings`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: use the std APIs the 1.91 MSRV already pays for

The crate declares rust-version 1.91.1, so a few things the code still
worked around are plain std now. All of them come from the 1.85–1.91
release notes; nothing here needs a newer toolchain than the manifest
already requires.

- std::sync::LazyLock (1.80) replaces the lazy_static! block in
  metrics.rs, and the lazy_static dependency goes. Every use site reads
  the same through Deref, so no caller changes.
- Duration::from_mins / from_hours (1.91) replace `from_secs(v * 60)`
  and `from_secs(v * 3600)` in the option parser and the shard-location
  refresh TTLs. One difference for the parser: an absurd count that
  overflows u64 seconds now panics in release builds too, where the
  multiplication used to wrap.
- Result::flatten (1.89) replaces `.and_then(|r| r)` on the replication
  join handle.
- OsStr::display (1.87) replaces `to_string_lossy()` where the name was
  only being formatted; the output is byte-identical.
- `#[allow]` becomes `#[expect]` (1.81) on the suppressions that are
  meant to be permanent, so a suppression that stops being needed
  becomes a warning rather than lingering. Doing that found four that
  already had: dead_code on ChunkManifest, base_name and last_io_error,
  and too_many_arguments on read_from_data_shards, which is down to
  seven parameters. Those attributes are deleted. The three allows that
  depend on cfg (a unix-only mutation, a linux-only field set, a
  profiling-only parameter) stay as allow, because expect would be
  unfulfilled on the other platforms.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* ci: add a commented-out clippy step to both Rust workflows

Both crates are warning-free under `cargo clippy --all-targets
-D warnings` now. Whether that becomes a gate is a policy call, so the
step is present but commented out; uncommenting it is the whole change.
The comment points at the `[lints.clippy]` table where crate-wide
exceptions are recorded, so the gate does not become a reason to
sprinkle allows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: guard parse_duration against overflow panics

Duration::from_mins/from_hours panic when the count overflows u64
seconds. Use checked_mul so an oversized CLI value falls back to the
parser default instead of crashing volume startup.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
This commit is contained in:
Eliah Rusin
2026-09-14 11:29:29 -07:00
committed by GitHub
co-authored by Claude Fable 5.1 Chris Lu
parent 49f20489e4
commit adaf3534fa
40 changed files with 1610 additions and 1579 deletions
@@ -59,6 +59,12 @@ jobs:
- name: Build Rust volume server
run: cd seaweed-volume && cargo build --release
# The crate is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[lints.clippy]` in
# seaweed-volume/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-volume && cargo clippy --all-targets -- -D warnings
- name: Run Rust unit tests
run: cd seaweed-volume && cargo test
+6
View File
@@ -73,6 +73,12 @@ jobs:
- name: Build the plugin workers
run: cd seaweed-worker && cargo build --release
# The workspace is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[workspace.lints.clippy]`
# in seaweed-worker/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-worker && cargo clippy --workspace --all-targets -- -D warnings
# The tests that need a live gateway skip themselves without one, the way
# the Go integration tests skip without Docker; the lifecycle suite in
# test/s3tables/lifecycle is what runs them against a real cluster.
-1
View File
@@ -4582,7 +4582,6 @@ dependencies = [
"image",
"jsonwebtoken",
"kamadak-exif",
"lazy_static",
"libc",
"md-5",
"memmap2",
+8 -1
View File
@@ -23,6 +23,14 @@ default = ["5bytes"]
# Pulls redb's experimental_cursor (and therefore experimental-api-5).
redb-experimental-cursor = ["redb/experimental_cursor"]
[lints.clippy]
# Every RPC path returns tonic::Status (176 bytes). Boxing it would change
# every handler signature for no gain, so the large-Err lint is off.
result_large_err = "allow"
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
[dependencies]
# Async runtime
tokio = { version = "1", features = ["full"] }
@@ -48,7 +56,6 @@ clap = { version = "4", features = ["derive"] }
# Metrics
prometheus = { version = "0.13", default-features = false, features = ["process"] }
lazy_static = "1"
# JWT
jsonwebtoken = { version = "10", features = ["rust_crypto"] }
+80 -73
View File
@@ -371,17 +371,18 @@ fn merge_options_file(args: Vec<String>) -> Vec<String> {
if arg == "--" {
break;
}
if arg.starts_with("--") {
let key = if let Some(eq) = arg.find('=') {
arg[2..eq].to_string()
if let Some(long) = arg.strip_prefix("--") {
let key = if let Some(eq) = long.find('=') {
long[..eq].to_string()
} else {
arg[2..].to_string()
long.to_string()
};
cli_flags.insert(key);
} else if arg.starts_with('-') && arg.len() > 2 {
} else if arg.len() > 2
&& let Some(without_dash) = arg.strip_prefix('-')
{
// Single-dash long option (already normalized to -- at this point,
// but handle both for safety)
let without_dash = &arg[1..];
let key = if let Some(eq) = without_dash.find('=') {
without_dash[..eq].to_string()
} else {
@@ -401,15 +402,14 @@ fn merge_options_file(args: Vec<String>) -> Vec<String> {
}
// Split on first `=`, ` `, or `:`
let (name, value) =
if let Some(pos) = trimmed.find(|c: char| c == '=' || c == ' ' || c == ':') {
(
trimmed[..pos].trim().to_string(),
trimmed[pos + 1..].trim().to_string(),
)
} else {
(trimmed.to_string(), String::new())
};
let (name, value) = if let Some(pos) = trimmed.find(['=', ' ', ':']) {
(
trimmed[..pos].trim().to_string(),
trimmed[pos + 1..].trim().to_string(),
)
} else {
(trimmed.to_string(), String::new())
};
// Strip leading dashes from name
let name = name.trim_start_matches('-').to_string();
@@ -436,10 +436,8 @@ fn merge_options_file(args: Vec<String>) -> Vec<String> {
/// Extract the options file path from args (looks for --options or -options).
fn find_options_arg(args: &[String]) -> String {
for i in 1..args.len() {
if args[i] == "--options" || args[i] == "-options" {
if i + 1 < args.len() {
return args[i + 1].clone();
}
if (args[i] == "--options" || args[i] == "-options") && i + 1 < args.len() {
return args[i + 1].clone();
}
if let Some(rest) = args[i].strip_prefix("--options=") {
return rest.to_string();
@@ -457,20 +455,22 @@ fn parse_duration(s: &str) -> std::time::Duration {
if s.is_empty() {
return std::time::Duration::from_secs(60);
}
if let Some(secs) = s.strip_suffix('s') {
if let Ok(v) = secs.parse::<u64>() {
return std::time::Duration::from_secs(v);
}
if let Some(secs) = s.strip_suffix('s')
&& let Ok(v) = secs.parse::<u64>()
{
return std::time::Duration::from_secs(v);
}
if let Some(mins) = s.strip_suffix('m') {
if let Ok(v) = mins.parse::<u64>() {
return std::time::Duration::from_secs(v * 60);
}
if let Some(mins) = s.strip_suffix('m')
&& let Ok(v) = mins.parse::<u64>()
&& let Some(seconds) = v.checked_mul(60)
{
return std::time::Duration::from_secs(seconds);
}
if let Some(hours) = s.strip_suffix('h') {
if let Ok(v) = hours.parse::<u64>() {
return std::time::Duration::from_secs(v * 3600);
}
if let Some(hours) = s.strip_suffix('h')
&& let Ok(v) = hours.parse::<u64>()
&& let Some(seconds) = v.checked_mul(3600)
{
return std::time::Duration::from_secs(seconds);
}
// Fallback: try parsing as raw seconds
if let Ok(v) = s.parse::<u64>() {
@@ -503,40 +503,40 @@ fn parse_min_free_spaces(min_free_space: &str, min_free_space_percent: &str) ->
}
// Try parsing human-readable bytes: e.g. "10GiB", "500MiB", "1TiB"
let s_upper = s.to_uppercase();
if let Some(rest) = s_upper.strip_suffix("TIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("TIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("KIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("KIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("TB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1_000_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("TB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1_000_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1_000_000.0) as u64);
}
// Default: 1%
MinFreeSpace::Percent(1.0)
@@ -1028,20 +1028,20 @@ pub fn parse_security_config(path: &str) -> SecurityConfig {
"cipher_suites" => cfg.tls_policy.cipher_suites = value.to_string(),
_ => {}
},
Section::Guard => match key {
"white_list" => {
Section::Guard => {
if key == "white_list" {
cfg.guard_white_list = value
.split(',')
.map(|s| s.trim().to_string())
.filter(|s| !s.is_empty())
.collect();
}
_ => {}
},
Section::Access => match key {
"ui" => cfg.access_ui = value.parse().unwrap_or(false),
_ => {}
},
}
Section::Access => {
if key == "ui" {
cfg.access_ui = value.parse().unwrap_or(false)
}
}
Section::None => {}
}
}
@@ -1188,12 +1188,11 @@ fn apply_env_overrides(cfg: &mut SecurityConfig) {
/// Mirrors Go's `util.DetectedHostAddress()`.
fn detect_host_address() -> String {
// Connect to a remote address to determine the local outbound IP
if let Ok(socket) = UdpSocket::bind("0.0.0.0:0") {
if socket.connect("8.8.8.8:80").is_ok() {
if let Ok(addr) = socket.local_addr() {
return addr.ip().to_string();
}
}
if let Ok(socket) = UdpSocket::bind("0.0.0.0:0")
&& socket.connect("8.8.8.8:80").is_ok()
&& let Ok(addr) = socket.local_addr()
{
return addr.ip().to_string();
}
"localhost".to_string()
}
@@ -1297,6 +1296,14 @@ mod tests {
assert_eq!(parse_duration("1h"), std::time::Duration::from_secs(3600));
assert_eq!(parse_duration("30"), std::time::Duration::from_secs(30));
assert_eq!(parse_duration(""), std::time::Duration::from_secs(60));
assert_eq!(
parse_duration("307445734561825861m"),
std::time::Duration::from_secs(60)
);
assert_eq!(
parse_duration("5124095576030432h"),
std::time::Duration::from_secs(60)
);
}
#[test]
+220 -122
View File
@@ -6,7 +6,7 @@ use prometheus::{
self, Encoder, GaugeVec, HistogramOpts, HistogramVec, IntCounter, IntCounterVec, IntGauge,
IntGaugeVec, Opts, Registry, TextEncoder,
};
use std::sync::Once;
use std::sync::{LazyLock, Once};
use crate::version;
@@ -16,220 +16,320 @@ pub struct PushGatewayConfig {
pub interval_seconds: u32,
}
lazy_static::lazy_static! {
pub static ref REGISTRY: Registry = Registry::new();
pub static REGISTRY: LazyLock<Registry> = LazyLock::new(Registry::new);
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static ref REQUEST_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_request_total", "Volume server requests"),
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static REQUEST_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_request_total",
"Volume server requests",
),
&["type", "code"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Request duration histogram with label `type` (HTTP method).
pub static ref REQUEST_DURATION: HistogramVec = HistogramVec::new(
/// Request duration histogram with label `type` (HTTP method).
pub static REQUEST_DURATION: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_request_seconds",
"Volume server request duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
/// Handler-level operation counter with label `type`.
pub static ref HANDLER_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_handler_total", "Volume server handler counters"),
/// Handler-level operation counter with label `type`.
pub static HANDLER_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_handler_total",
"Volume server handler counters",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
/// Vacuuming compact counter with label `success` (true/false).
pub static ref VACUUMING_COMPACT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_compact_count", "Counter of volume vacuuming Compact counter"),
/// Vacuuming compact counter with label `success` (true/false).
pub static VACUUMING_COMPACT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_compact_count",
"Counter of volume vacuuming Compact counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming commit counter with label `success` (true/false).
pub static ref VACUUMING_COMMIT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_commit_count", "Counter of volume vacuuming commit counter"),
/// Vacuuming commit counter with label `success` (true/false).
pub static VACUUMING_COMMIT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_commit_count",
"Counter of volume vacuuming commit counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static ref VACUUMING_HISTOGRAM: HistogramVec = HistogramVec::new(
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static VACUUMING_HISTOGRAM: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_vacuuming_seconds",
"Volume vacuuming duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
/// Volumes per collection and type (volume/ec_shards).
pub static ref VOLUME_GAUGE: GaugeVec = GaugeVec::new(
/// Volumes per collection and type (volume/ec_shards).
pub static VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_volumes", "Number of volumes"),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Read-only volumes per collection and type.
pub static ref READ_ONLY_VOLUME_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_read_only_volumes", "Number of read-only volumes."),
/// Read-only volumes per collection and type.
pub static READ_ONLY_VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_read_only_volumes",
"Number of read-only volumes.",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Maximum number of volumes this server can hold.
pub static ref MAX_VOLUMES: IntGauge = IntGauge::new(
/// Maximum number of volumes this server can hold.
pub static MAX_VOLUMES: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_max_volumes",
"Maximum number of volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static ref DISK_SIZE_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_total_disk_size", "Actual disk size used by volumes"),
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static DISK_SIZE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_total_disk_size",
"Actual disk size used by volumes",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
/// Disk resource usage per directory and type (all/used/free/avail).
pub static ref RESOURCE_GAUGE: GaugeVec = GaugeVec::new(
/// Disk resource usage per directory and type (all/used/free/avail).
pub static RESOURCE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_resource", "Server resource usage"),
&["name", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
/// In-flight requests per HTTP method.
pub static ref INFLIGHT_REQUESTS_GAUGE: IntGaugeVec = IntGaugeVec::new(
Opts::new("SeaweedFS_volumeServer_in_flight_requests", "Current number of in-flight requests being handled by volume server."),
/// In-flight requests per HTTP method.
pub static INFLIGHT_REQUESTS_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_in_flight_requests",
"Current number of in-flight requests being handled by volume server.",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent download limit in bytes.
pub static ref CONCURRENT_DOWNLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent download limit in bytes.
pub static CONCURRENT_DOWNLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_download_limit",
"Limit for total concurrent download size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent upload limit in bytes.
pub static ref CONCURRENT_UPLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent upload limit in bytes.
pub static CONCURRENT_UPLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_upload_limit",
"Limit for total concurrent upload size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight download bytes.
pub static ref INFLIGHT_DOWNLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight download bytes.
pub static INFLIGHT_DOWNLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_download_size",
"In flight total download size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight upload bytes.
pub static ref INFLIGHT_UPLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight upload bytes.
pub static INFLIGHT_UPLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_upload_size",
"In flight total upload size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static ref UPLOAD_ERROR_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created");
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static UPLOAD_ERROR_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created")
});
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static ref SCRUB_LAST_TIME_SECONDS: GaugeVec = GaugeVec::new(
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static SCRUB_LAST_TIME_SECONDS: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_last_time_seconds",
"Last scrub execution time, as seconds since UNIX epoch.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_VOLUME_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static SCRUB_VOLUME_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_volume_failures",
"Counter of overall volumes with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_SHARD_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static SCRUB_SHARD_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_shard_failures",
"Counter of overall EC shards with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of storage read/write EIO errors on volumes and EC shards.
/// Mirrors Go's VolumeServerStorageIoErrorCounter.
pub static ref STORAGE_IO_ERROR_COUNTER: IntCounter = IntCounter::new(
/// Counter of storage read/write EIO errors on volumes and EC shards.
/// Mirrors Go's VolumeServerStorageIoErrorCounter.
pub static STORAGE_IO_ERROR_COUNTER: LazyLock<IntCounter> = LazyLock::new(|| {
IntCounter::new(
"SeaweedFS_volumeServer_storage_io_error_total",
"Counter of storage read/write EIO errors on volumes and EC shards.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Number of volumes quarantined due to storage IO errors.
/// Mirrors Go's VolumeServerIoQuarantineGauge.
pub static ref IO_QUARANTINE_GAUGE: IntGaugeVec = IntGaugeVec::new(
/// Number of volumes quarantined due to storage IO errors.
/// Mirrors Go's VolumeServerIoQuarantineGauge.
pub static IO_QUARANTINE_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_io_quarantine",
"Number of volumes or EC shards quarantined due to storage IO errors.",
),
&["kind"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Legacy aliases for backward compat with existing code ----
// ---- Legacy aliases for backward compat with existing code ----
/// Total number of volumes on this server (flat gauge).
pub static ref VOLUMES_TOTAL: IntGauge = IntGauge::new(
"volume_server_volumes_total",
"Total number of volumes",
).expect("metric can be created");
/// Total number of volumes on this server (flat gauge).
pub static VOLUMES_TOTAL: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new("volume_server_volumes_total", "Total number of volumes")
.expect("metric can be created")
});
/// Disk size in bytes per directory.
pub static ref DISK_SIZE_BYTES: IntGaugeVec = IntGaugeVec::new(
/// Disk size in bytes per directory.
pub static DISK_SIZE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_size_bytes", "Disk size in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Disk free bytes per directory.
pub static ref DISK_FREE_BYTES: IntGaugeVec = IntGaugeVec::new(
/// Disk free bytes per directory.
pub static DISK_FREE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_free_bytes", "Disk free space in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current number of in-flight requests (flat gauge).
pub static ref INFLIGHT_REQUESTS: IntGauge = IntGauge::new(
/// Current number of in-flight requests (flat gauge).
pub static INFLIGHT_REQUESTS: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_inflight_requests",
"Current number of in-flight requests",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Total number of files stored across all volumes.
pub static ref VOLUME_FILE_COUNT: IntGauge = IntGauge::new(
/// Total number of files stored across all volumes.
pub static VOLUME_FILE_COUNT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_volume_file_count",
"Total number of files stored across all volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Build info (Go: BuildInfo) ----
// ---- Build info (Go: BuildInfo) ----
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static ref BUILD_INFO: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created");
}
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static BUILD_INFO: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created")
});
/// Generate exponential bucket boundaries for histograms.
fn exponential_buckets(start: f64, factor: f64, count: usize) -> Vec<f64> {
@@ -377,10 +477,8 @@ fn delete_partial_match_collection(gauge: &GaugeVec, collection: &str) {
type_value = Some(label.get_value().to_string());
}
}
if matches_collection {
if let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
if matches_collection && let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
}
}
@@ -173,10 +173,10 @@ pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool)
// same host wherever the matching relay exists (common in IPv6-only cloud).
// to_ipv4_mapped above only covers ::ffff: mapped addresses, so pull the
// embedded IPv4 out of the other forms and re-check it against the rules.
if let IpAddr::V6(v6) = ip {
if let Some(v4) = embedded_transition_ipv4(v6) {
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
if let IpAddr::V6(v6) = ip
&& let Some(v4) = embedded_transition_ipv4(v6)
{
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
Ok(())
}
@@ -214,9 +214,7 @@ fn precheck_endpoint(endpoint: &str) -> Result<HostCheck, String> {
// Authority is everything up to the first '/', '?', or '#'.
let after = &trimmed[scheme_end + 3..];
let authority_end = after
.find(|c| c == '/' || c == '?' || c == '#')
.unwrap_or(after.len());
let authority_end = after.find(['/', '?', '#']).unwrap_or(after.len());
let authority = &after[..authority_end];
// Strip optional userinfo ("user:pass@").
+4 -4
View File
@@ -297,10 +297,10 @@ impl Guard {
/// Extract host from "host:port" or "[::1]:port" format.
fn extract_host(addr: &str) -> String {
// Handle IPv6 with brackets
if addr.starts_with('[') {
if let Some(end) = addr.find(']') {
return addr[1..end].to_string();
}
if addr.starts_with('[')
&& let Some(end) = addr.find(']')
{
return addr[1..end].to_string();
}
// Handle host:port
if let Some(pos) = addr.rfind(':') {
+150 -148
View File
@@ -1236,7 +1236,7 @@ impl VolumeServer for VolumeGrpcService {
let mut store = self.state.store.write().unwrap();
store
.delete_collection(collection)
.map_err(|e| Status::internal(e))?;
.map_err(Status::internal)?;
}
// The delta the notify path derives is the only thing that tells the
// master these slots came free: a heartbeat carries the whole list only
@@ -1261,7 +1261,7 @@ impl VolumeServer for VolumeGrpcService {
} else {
Some(
crate::storage::needle::ttl::TTL::read(&req.ttl)
.map_err(|e| Status::invalid_argument(e))?,
.map_err(Status::invalid_argument)?,
)
};
let disk_type = DiskType::from_string(&req.disk_type);
@@ -2053,13 +2053,13 @@ impl VolumeServer for VolumeGrpcService {
// Remove the .note file. A leftover note fails the load on the
// next restart, so a removal failure must fail the copy.
if let Err(e) = std::fs::remove_file(&note_path) {
if e.kind() != std::io::ErrorKind::NotFound {
return Err(Status::internal(format!(
"remove .note for volume {}: {}",
vid, e
)));
}
if let Err(e) = std::fs::remove_file(&note_path)
&& e.kind() != std::io::ErrorKind::NotFound
{
return Err(Status::internal(format!(
"remove .note for volume {}: {}",
vid, e
)));
}
// Verify file sizes
@@ -2115,7 +2115,7 @@ impl VolumeServer for VolumeGrpcService {
// unmounts and deletes the replica it just created.
if tx
.send(Ok(volume_server_pb::VolumeCopyResponse {
last_append_at_ns: last_append_at_ns,
last_append_at_ns,
processed_bytes: 0,
}))
.await
@@ -3108,8 +3108,8 @@ impl VolumeServer for VolumeGrpcService {
dat_file_size,
expire_at_sec,
ec_shard_config: Some(crate::storage::volume::VifEcShardConfig {
data_shards: data_shards,
parity_shards: parity_shards,
data_shards,
parity_shards,
// This run's identity; the read path rejects a shard from a
// different encode run.
encode_ts_ns: std::time::SystemTime::now()
@@ -3833,15 +3833,14 @@ impl VolumeServer for VolumeGrpcService {
if let Some((_offset, size)) = ec_vol
.find_needle_from_ecx(needle_id)
.map_err(|e| Status::internal(e.to_string()))?
&& size.is_deleted()
{
if size.is_deleted() {
let results = vec![Ok(volume_server_pb::VolumeEcShardReadResponse {
is_deleted: true,
encode_ts_ns: served_encode_ts_ns,
..Default::default()
})];
return Ok(Response::new(Box::pin(tokio_stream::iter(results))));
}
let results = vec![Ok(volume_server_pb::VolumeEcShardReadResponse {
is_deleted: true,
encode_ts_ns: served_encode_ts_ns,
..Default::default()
})];
return Ok(Response::new(Box::pin(tokio_stream::iter(results))));
}
}
@@ -3915,13 +3914,13 @@ impl VolumeServer for VolumeGrpcService {
let mut store = self.state.store.write().unwrap();
if let Some(ec_vol) = store.find_ec_volume_mut(vid) {
// Check if already deleted via ecx index
if let Ok(Some((_offset, size))) = ec_vol.find_needle_from_ecx(needle_id) {
if size.is_deleted() {
// Already deleted, no-op
return Ok(Response::new(
volume_server_pb::VolumeEcBlobDeleteResponse {},
));
}
if let Ok(Some((_offset, size))) = ec_vol.find_needle_from_ecx(needle_id)
&& size.is_deleted()
{
// Already deleted, no-op
return Ok(Response::new(
volume_server_pb::VolumeEcBlobDeleteResponse {},
));
}
ec_vol
.journal_delete(needle_id)
@@ -4056,8 +4055,8 @@ impl VolumeServer for VolumeGrpcService {
}
// Check that all data shards are present somewhere on this server.
for shard_id in 0..data_shards {
if shard_dirs[shard_id].is_none() {
for (shard_id, dir) in shard_dirs[..data_shards].iter().enumerate() {
if dir.is_none() {
return Err(Status::internal(format!(
"ec volume {} missing shard {}",
req.volume_id, shard_id
@@ -4560,15 +4559,15 @@ impl VolumeServer for VolumeGrpcService {
// a local .dat exists, so leaving one wedges every
// retry on "already on local disk", and a restart would
// load the sparse file as the volume's data.
if let Err(rm) = std::fs::remove_file(&dat_path) {
if rm.kind() != std::io::ErrorKind::NotFound {
tracing::warn!(
"volume {} could not remove the incomplete download {}: {}",
vid,
dat_path,
rm
);
}
if let Err(rm) = std::fs::remove_file(&dat_path)
&& rm.kind() != std::io::ErrorKind::NotFound
{
tracing::warn!(
"volume {} could not remove the incomplete download {}: {}",
vid,
dat_path,
rm
);
}
Status::internal(format!(
"backend {} copy file {}: {}",
@@ -4823,16 +4822,13 @@ impl VolumeServer for VolumeGrpcService {
// provider default (e.g. real AWS S3) and cannot target an internal
// host, so skip it. Extends the Go volume server's validateRemoteEndpoint
// gate, which only covered type "s3".
if !self.state.allow_untrusted_remote_endpoints {
if let Some(endpoint) = crate::remote_storage::s3_compatible_endpoint(remote_conf) {
if !endpoint.trim().is_empty() {
crate::remote_storage::validate_remote_endpoint(endpoint)
.await
.map_err(|e| {
Status::invalid_argument(format!("reject remote endpoint: {}", e))
})?;
}
}
if !self.state.allow_untrusted_remote_endpoints
&& let Some(endpoint) = crate::remote_storage::s3_compatible_endpoint(remote_conf)
&& !endpoint.trim().is_empty()
{
crate::remote_storage::validate_remote_endpoint(endpoint)
.await
.map_err(|e| Status::invalid_argument(format!("reject remote endpoint: {}", e)))?;
}
// Create remote storage client
@@ -5003,7 +4999,7 @@ impl VolumeServer for VolumeGrpcService {
// Validate mode
let mode = req.mode;
match mode {
1 | 2 | 3 | 4 | 5 => {} // INDEX=1, FULL=2, LOCAL=3, CHECKSUM=4, READS=5
1..=5 => {} // INDEX=1, FULL=2, LOCAL=3, CHECKSUM=4, READS=5
_ => {
return Err(Status::invalid_argument(format!(
"unsupported EC volume scrub mode {}",
@@ -5053,7 +5049,7 @@ impl VolumeServer for VolumeGrpcService {
let mut stripes: Vec<Result<volume_server_pb::QueriedStripe, Status>> = Vec::new();
for fid_str in &req.from_file_ids {
let file_id = needle::FileId::parse(fid_str).map_err(|e| Status::internal(e))?;
let file_id = needle::FileId::parse(fid_str).map_err(Status::internal)?;
let mut n = Needle {
id: file_id.key,
@@ -5081,13 +5077,13 @@ impl VolumeServer for VolumeGrpcService {
let input = req.input_serialization.as_ref();
// CSV input: no output (Go does nothing for CSV)
if input.map_or(false, |i| i.csv_input.is_some()) {
if input.is_some_and(|i| i.csv_input.is_some()) {
// No stripes emitted for CSV
continue;
}
// JSON input: process lines
if input.map_or(false, |i| i.json_input.is_some()) {
if input.is_some_and(|i| i.json_input.is_some()) {
let filter = req.filter.as_ref();
let data_str = String::from_utf8_lossy(&n.data);
let mut records: Vec<u8> = Vec::new();
@@ -5102,69 +5098,70 @@ impl VolumeServer for VolumeGrpcService {
};
// Apply filter
if let Some(f) = filter {
if !f.field.is_empty() && !f.operand.is_empty() {
let field_val = &parsed[&f.field];
let pass = match f.operand.as_str() {
">" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv > tv
} else {
false
}
if let Some(f) = filter
&& !f.field.is_empty()
&& !f.operand.is_empty()
{
let field_val = &parsed[&f.field];
let pass = match f.operand.as_str() {
">" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv > tv
} else {
false
}
">=" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv >= tv
} else {
false
}
}
"<" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv < tv
} else {
false
}
}
"<=" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv <= tv
} else {
false
}
}
"=" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv == tv
} else {
field_val.as_str().map_or(false, |s| s == f.value)
}
}
"!=" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv != tv
} else {
field_val.as_str().map_or(true, |s| s != f.value)
}
}
_ => true,
};
if !pass {
continue;
}
">=" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv >= tv
} else {
false
}
}
"<" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv < tv
} else {
false
}
}
"<=" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv <= tv
} else {
false
}
}
"=" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv == tv
} else {
field_val.as_str().is_some_and(|s| s == f.value)
}
}
"!=" => {
if let (Some(fv), Ok(tv)) =
(field_val.as_f64(), f.value.parse::<f64>())
{
fv != tv
} else {
field_val.as_str().is_none_or(|s| s != f.value)
}
}
_ => true,
};
if !pass {
continue;
}
}
@@ -5208,7 +5205,7 @@ impl VolumeServer for VolumeGrpcService {
let store = self.state.store.read().unwrap();
// Try normal volume first
if let Some(_) = store.find_volume(vid) {
if store.find_volume(vid).is_some() {
let mut n = Needle {
id: needle_id,
..Needle::default()
@@ -5503,6 +5500,7 @@ async fn drain_copy_stream_to_file(
/// Copy a file from a remote volume server via CopyFile streaming RPC.
/// Returns the modified_ts_ns received from the source.
#[expect(clippy::too_many_arguments)]
async fn copy_file_from_source<T>(
client: &mut volume_server_pb::volume_server_client::VolumeServerClient<T>,
is_ec_volume: bool,
@@ -5959,41 +5957,40 @@ mod tests {
if let Some(range) = headers
.get(header::RANGE)
.and_then(|value| value.to_str().ok())
&& let Some(range_value) = range.strip_prefix("bytes=")
{
if let Some(range_value) = range.strip_prefix("bytes=") {
let mut parts = range_value.splitn(2, '-');
let start = parts
.next()
.and_then(|value| value.parse::<usize>().ok())
.unwrap_or(0);
let end = parts
.next()
.and_then(|value| value.parse::<usize>().ok())
.unwrap_or_else(|| bytes.len().saturating_sub(1));
let start = start.min(bytes.len());
let end = end.min(bytes.len().saturating_sub(1));
let payload = if start > end || start >= bytes.len() {
Vec::new()
} else {
bytes[start..=end].to_vec()
};
let mut response_headers = HeaderMap::new();
response_headers.insert(
header::CONTENT_RANGE,
HeaderValue::from_str(&format!(
"bytes {}-{}/{}",
start,
end,
bytes.len()
))
.unwrap(),
);
response_headers.insert(
header::CONTENT_LENGTH,
HeaderValue::from_str(&payload.len().to_string()).unwrap(),
);
return (StatusCode::PARTIAL_CONTENT, response_headers, payload);
}
let mut parts = range_value.splitn(2, '-');
let start = parts
.next()
.and_then(|value| value.parse::<usize>().ok())
.unwrap_or(0);
let end = parts
.next()
.and_then(|value| value.parse::<usize>().ok())
.unwrap_or_else(|| bytes.len().saturating_sub(1));
let start = start.min(bytes.len());
let end = end.min(bytes.len().saturating_sub(1));
let payload = if start > end || start >= bytes.len() {
Vec::new()
} else {
bytes[start..=end].to_vec()
};
let mut response_headers = HeaderMap::new();
response_headers.insert(
header::CONTENT_RANGE,
HeaderValue::from_str(&format!(
"bytes {}-{}/{}",
start,
end,
bytes.len()
))
.unwrap(),
);
response_headers.insert(
header::CONTENT_LENGTH,
HeaderValue::from_str(&payload.len().to_string()).unwrap(),
);
return (StatusCode::PARTIAL_CONTENT, response_headers, payload);
}
let mut response_headers = HeaderMap::new();
@@ -6734,6 +6731,11 @@ mod tests {
// very first check and returns before mount_volume, exercising the wrong
// path — the test would be green for the wrong reason.
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
#[expect(
clippy::await_holding_lock,
clippy::readonly_write_lock,
reason = "the store write guard is a barrier that parks the copy task at the mount block"
)]
async fn test_volume_copy_after_mount_cancellation_rolls_back_mount() {
let (source_service, _source_tmp, _dat_bytes) = make_local_service_with_large_volume();
let (port, _shutdown) = serve_source(source_service).await;
+252 -275
View File
@@ -193,10 +193,7 @@ impl http_body::Body for StreamingBody {
}
Ok(Err(e)) => return std::task::Poll::Ready(Some(Err(e))),
Err(e) => {
return std::task::Poll::Ready(Some(Err(std::io::Error::new(
std::io::ErrorKind::Other,
e,
))))
return std::task::Poll::Ready(Some(Err(std::io::Error::other(e))));
}
}
}
@@ -322,10 +319,9 @@ fn parse_url_path(path: &str) -> Option<(VolumeId, NeedleId, Cookie)> {
// Try "vid,fid" or "vid/fid" or "vid/fid/filename" formats
let (vid_str, fid_part) = if let Some(pos) = path.find(',') {
(&path[..pos], &path[pos + 1..])
} else if let Some(pos) = path.find('/') {
(&path[..pos], &path[pos + 1..])
} else {
return None;
let pos = path.find('/')?;
(&path[..pos], &path[pos + 1..])
};
// For fid part, strip extension from the fid (not from filename)
@@ -409,10 +405,10 @@ async fn lookup_volume(
.json()
.await
.map_err(|e| format!("lookup parse failed: {}", e))?;
if let Some(err) = result.error {
if !err.is_empty() {
return Err(err);
}
if let Some(err) = result.error
&& !err.is_empty()
{
return Err(err);
}
Ok(result.locations.unwrap_or_default())
}
@@ -688,7 +684,8 @@ fn build_proxy_request_info(
raw_fid
};
(trimmed[..pos].to_string(), fid.to_string())
} else if let Some(pos) = trimmed.find('/') {
} else {
let pos = trimmed.find('/')?;
let after = &trimmed[pos + 1..];
let fid_part = if let Some(slash) = after.find('/') {
&after[..slash]
@@ -696,8 +693,6 @@ fn build_proxy_request_info(
after
};
(trimmed[..pos].to_string(), fid_part.to_string())
} else {
return None;
};
Some(ProxyRequestInfo {
@@ -837,12 +832,12 @@ fn redirect_request(info: &ProxyRequestInfo, target: &VolumeLocation, scheme: &s
let mut query_params = Vec::new();
if !info.original_query.is_empty() {
for param in info.original_query.split('&') {
if let Some((key, value)) = param.split_once('=') {
if key == "collection" {
query_params.push(format!("collection={}", value));
}
// Intentionally drop readDeleted and other params (Go parity)
if let Some((key, value)) = param.split_once('=')
&& key == "collection"
{
query_params.push(format!("collection={}", value));
}
// Intentionally drop readDeleted and other params (Go parity)
}
}
query_params.push("proxied=true".to_string());
@@ -852,7 +847,7 @@ fn redirect_request(info: &ProxyRequestInfo, target: &VolumeLocation, scheme: &s
let target_http = to_http_address(&target.url);
let raw_target = format!(
"{}/{},{}?{}",
target_http, &info.vid_str, &info.fid_str, query
target_http, info.vid_str, info.fid_str, query
);
let location = match normalize_outgoing_http_url(scheme, &raw_target) {
Ok(url) => url,
@@ -954,12 +949,12 @@ async fn get_or_head_handler_inner(
// so invalid paths with JWT enabled return 401, not 400.
let file_id = extract_file_id(&path);
let token = extract_jwt(&headers, request.uri());
if let Err(_) =
state
.guard
.read()
.unwrap()
.check_jwt_for_file(token.as_deref(), &file_id, false)
if state
.guard
.read()
.unwrap()
.check_jwt_for_file(token.as_deref(), &file_id, false)
.is_err()
{
let body = serde_json::json!({"error": "wrong jwt"});
return Response::builder()
@@ -1015,16 +1010,15 @@ async fn get_or_head_handler_inner(
let should_try_replica =
!query_string.contains("proxied=true") && !state.master_url.is_empty() && {
let store = state.store.read().unwrap();
store.find_volume(vid).map_or(false, |(_, vol)| {
store.find_volume(vid).is_some_and(|(_, vol)| {
vol.super_block.replica_placement.get_copy_count() > 1
})
};
if should_try_replica {
if let Some(info) =
if should_try_replica
&& let Some(info) =
build_proxy_request_info(&path, request.headers(), &query_string)
{
return proxy_or_redirect_to_target(&state, info, vid, true).await;
}
{
return proxy_or_redirect_to_target(&state, info, vid, true).await;
}
// Blocking wait loop (Go's waitForDownloadSlot)
@@ -1231,57 +1225,53 @@ async fn get_or_head_handler_inner(
// Build Last-Modified header (RFC 1123 format) — must be done before conditional checks
let last_modified_str = if n.last_modified > 0 {
use chrono::{TimeZone, Utc};
if let Some(dt) = Utc.timestamp_opt(n.last_modified as i64, 0).single() {
Some(dt.format("%a, %d %b %Y %H:%M:%S GMT").to_string())
} else {
None
}
Utc.timestamp_opt(n.last_modified as i64, 0)
.single()
.map(|dt| dt.format("%a, %d %b %Y %H:%M:%S GMT").to_string())
} else {
None
};
// Check If-Modified-Since FIRST (Go checks this before If-None-Match)
if n.last_modified > 0 {
if let Some(ims_header) = headers.get(header::IF_MODIFIED_SINCE) {
if let Ok(ims_str) = ims_header.to_str() {
// Parse HTTP date format: "Mon, 02 Jan 2006 15:04:05 GMT"
if let Ok(ims_time) =
chrono::NaiveDateTime::parse_from_str(ims_str, "%a, %d %b %Y %H:%M:%S GMT")
{
if (n.last_modified as i64) <= ims_time.and_utc().timestamp() {
let mut resp = StatusCode::NOT_MODIFIED.into_response();
if let Some(ref lm) = last_modified_str {
resp.headers_mut()
.insert(header::LAST_MODIFIED, lm.parse().unwrap());
}
// Go sets ETag AFTER the 304 return paths (L235), so 304 does NOT include ETag
return resp;
}
}
if n.last_modified > 0
&& let Some(ims_header) = headers.get(header::IF_MODIFIED_SINCE)
&& let Ok(ims_str) = ims_header.to_str()
{
// Parse HTTP date format: "Mon, 02 Jan 2006 15:04:05 GMT"
if let Ok(ims_time) =
chrono::NaiveDateTime::parse_from_str(ims_str, "%a, %d %b %Y %H:%M:%S GMT")
&& (n.last_modified as i64) <= ims_time.and_utc().timestamp()
{
let mut resp = StatusCode::NOT_MODIFIED.into_response();
if let Some(ref lm) = last_modified_str {
resp.headers_mut()
.insert(header::LAST_MODIFIED, lm.parse().unwrap());
}
// Go sets ETag AFTER the 304 return paths (L235), so 304 does NOT include ETag
return resp;
}
}
// Check If-None-Match SECOND
if let Some(if_none_match) = headers.get(header::IF_NONE_MATCH) {
if let Ok(inm) = if_none_match.to_str() {
if inm == etag {
let mut resp = StatusCode::NOT_MODIFIED.into_response();
if let Some(ref lm) = last_modified_str {
resp.headers_mut()
.insert(header::LAST_MODIFIED, lm.parse().unwrap());
}
// Go sets ETag AFTER the 304 return paths (L235), so 304 does NOT include ETag
return resp;
}
if let Some(if_none_match) = headers.get(header::IF_NONE_MATCH)
&& let Ok(inm) = if_none_match.to_str()
&& inm == etag
{
let mut resp = StatusCode::NOT_MODIFIED.into_response();
if let Some(ref lm) = last_modified_str {
resp.headers_mut()
.insert(header::LAST_MODIFIED, lm.parse().unwrap());
}
// Go sets ETag AFTER the 304 return paths (L235), so 304 does NOT include ETag
return resp;
}
// Chunk manifest expansion (needs full data) — after conditional checks, before response
// Pass ETag so chunk manifest responses include it (matches Go: ETag is set on the
// response writer before tryHandleChunkedFile runs).
if n.is_chunk_manifest() && !bypass_cm {
if let Some(resp) = try_expand_chunk_manifest(
if n.is_chunk_manifest()
&& !bypass_cm
&& let Some(resp) = try_expand_chunk_manifest(
&state,
&n,
&headers,
@@ -1292,27 +1282,26 @@ async fn get_or_head_handler_inner(
&last_modified_str,
)
.await
{
return resp;
}
// If manifest expansion fails (invalid JSON etc.), fall through to raw data
{
return resp;
}
// If manifest expansion fails (invalid JSON etc.), fall through to raw data
let mut response_headers = HeaderMap::new();
response_headers.insert(header::ETAG, etag.parse().unwrap());
// H1: Emit pairs as response headers
if n.has_pairs() && !n.pairs.is_empty() {
if let Ok(pair_map) =
if n.has_pairs()
&& !n.pairs.is_empty()
&& let Ok(pair_map) =
serde_json::from_slice::<std::collections::HashMap<String, String>>(&n.pairs)
{
for (k, v) in &pair_map {
if let (Ok(hname), Ok(hval)) = (
axum::http::HeaderName::from_bytes(k.as_bytes()),
axum::http::HeaderValue::from_str(v),
) {
response_headers.insert(hname, hval);
}
{
for (k, v) in &pair_map {
if let (Ok(hname), Ok(hval)) = (
axum::http::HeaderName::from_bytes(k.as_bytes()),
axum::http::HeaderValue::from_str(v),
) {
response_headers.insert(hname, hval);
}
}
}
@@ -1322,10 +1311,10 @@ async fn get_or_head_handler_inner(
let mut ext = ext;
if n.name_size > 0 && filename.is_empty() {
filename = String::from_utf8_lossy(&n.name).to_string();
if ext.is_empty() {
if let Some(dot_pos) = filename.rfind('.') {
ext = filename[dot_pos..].to_lowercase();
}
if ext.is_empty()
&& let Some(dot_pos) = filename.rfind('.')
{
ext = filename[dot_pos..].to_lowercase();
}
}
@@ -1429,80 +1418,72 @@ async fn get_or_head_handler_inner(
}
// ---- Streaming path: large uncompressed files ----
if can_stream {
if let Some(info) = stream_info {
response_headers.insert(header::ACCEPT_RANGES, "bytes".parse().unwrap());
response_headers.insert(
header::CONTENT_LENGTH,
info.data_size.to_string().parse().unwrap(),
);
if can_stream && let Some(info) = stream_info {
response_headers.insert(header::ACCEPT_RANGES, "bytes".parse().unwrap());
response_headers.insert(
header::CONTENT_LENGTH,
info.data_size.to_string().parse().unwrap(),
);
let tracked_bytes = info.data_size as i64;
let tracking_state = if download_guard.is_some() {
let new_val = state
.inflight_download_bytes
.fetch_add(tracked_bytes, Ordering::Relaxed)
+ tracked_bytes;
metrics::INFLIGHT_DOWNLOAD_SIZE.set(new_val);
Some(state.clone())
} else {
let tracked_bytes = info.data_size as i64;
let tracking_state = if download_guard.is_some() {
let new_val = state
.inflight_download_bytes
.fetch_add(tracked_bytes, Ordering::Relaxed)
+ tracked_bytes;
metrics::INFLIGHT_DOWNLOAD_SIZE.set(new_val);
Some(state.clone())
} else {
None
};
let streaming = StreamingBody {
source: info.source,
data_offset: info.data_file_offset,
data_size: info.data_size,
pos: 0,
chunk_size: streaming_chunk_size(state.read_buffer_size_bytes, info.data_size as usize),
_held_read_lease: if state.has_slow_read {
None
};
} else {
Some(info.data_file_access_control.read_lock())
},
data_file_access_control: info.data_file_access_control,
hold_read_lock_for_stream: !state.has_slow_read,
pending: None,
state: tracking_state,
tracked_bytes,
server_state: state.clone(),
volume_id: info.volume_id,
needle_id: info.needle_id,
compaction_revision: info.compaction_revision,
};
let streaming = StreamingBody {
source: info.source,
data_offset: info.data_file_offset,
data_size: info.data_size,
pos: 0,
chunk_size: streaming_chunk_size(
state.read_buffer_size_bytes,
info.data_size as usize,
),
_held_read_lease: if state.has_slow_read {
None
} else {
Some(info.data_file_access_control.read_lock())
},
data_file_access_control: info.data_file_access_control,
hold_read_lock_for_stream: !state.has_slow_read,
pending: None,
state: tracking_state,
tracked_bytes,
server_state: state.clone(),
volume_id: info.volume_id,
needle_id: info.needle_id,
compaction_revision: info.compaction_revision,
};
let body = Body::new(streaming);
let mut resp = Response::new(body);
*resp.status_mut() = StatusCode::OK;
*resp.headers_mut() = response_headers;
return resp;
}
let body = Body::new(streaming);
let mut resp = Response::new(body);
*resp.status_mut() = StatusCode::OK;
*resp.headers_mut() = response_headers;
return resp;
}
if can_handle_head_from_meta {
if let Some(info) = stream_info {
response_headers.insert(
header::CONTENT_LENGTH,
info.data_size.to_string().parse().unwrap(),
);
return (StatusCode::OK, response_headers).into_response();
}
if can_handle_head_from_meta && let Some(info) = stream_info {
response_headers.insert(
header::CONTENT_LENGTH,
info.data_size.to_string().parse().unwrap(),
);
return (StatusCode::OK, response_headers).into_response();
}
if can_handle_range_from_source {
if let (Some(range_header), Some(info)) = (headers.get(header::RANGE), stream_info) {
if let Ok(range_str) = range_header.to_str() {
return handle_range_request_from_source(
range_str,
info,
response_headers,
track_download.then(|| state.clone()),
);
}
}
if can_handle_range_from_source
&& let (Some(range_header), Some(info)) = (headers.get(header::RANGE), stream_info)
&& let Ok(range_str) = range_header.to_str()
{
return handle_range_request_from_source(
range_str,
info,
response_headers,
track_download.then(|| state.clone()),
);
}
// ---- Buffered path: small files, compressed, images, range requests ----
@@ -1572,15 +1553,15 @@ async fn get_or_head_handler_inner(
response_headers.insert(header::ACCEPT_RANGES, "bytes".parse().unwrap());
// Check Range header
if let Some(range_header) = headers.get(header::RANGE) {
if let Ok(range_str) = range_header.to_str() {
return handle_range_request(
range_str,
&data,
response_headers,
track_download.then(|| state.clone()),
);
}
if let Some(range_header) = headers.get(header::RANGE)
&& let Ok(range_str) = range_header.to_str()
{
return handle_range_request(
range_str,
&data,
response_headers,
track_download.then(|| state.clone()),
);
}
if method == Method::HEAD {
@@ -2006,7 +1987,7 @@ fn extract_extension_from_path(path: &str) -> String {
if let Some(dot_pos) = filename.rfind('.') {
return filename[dot_pos..].to_lowercase();
}
} else if parts.len() >= 1 {
} else if !parts.is_empty() {
// 2-segment path: /vid,fid.ext or /vid/fid.ext
// Go's parseURLPath extracts ext from the full path for all formats
let last = parts[parts.len() - 1];
@@ -2129,7 +2110,7 @@ pub async fn post_handler(
// Go's r.ParseForm() returns 400 on malformed query strings
return json_error_with_query(
StatusCode::BAD_REQUEST,
&format!("form parse error: {}", e),
format!("form parse error: {}", e),
Some(&query),
);
}
@@ -2145,11 +2126,12 @@ pub async fn post_handler(
// JWT check for writes
let file_id = extract_file_id(&path);
let token = extract_jwt(&headers, request.uri());
if let Err(_) = state
if state
.guard
.read()
.unwrap()
.check_jwt_for_file(token.as_deref(), &file_id, true)
.is_err()
{
return json_error_with_query(StatusCode::UNAUTHORIZED, "wrong jwt", Some(&query));
}
@@ -2279,11 +2261,8 @@ pub async fn post_handler(
.split(';')
.find_map(|part| {
let part = part.trim();
if let Some(val) = part.strip_prefix("boundary=") {
Some(val.trim_matches('"').to_string())
} else {
None
}
part.strip_prefix("boundary=")
.map(|val| val.trim_matches('"').to_string())
})
.unwrap_or_default();
@@ -2425,17 +2404,17 @@ pub async fn post_handler(
} else {
None
};
if let (Some(expected_md5), Some(actual_md5)) = (&content_md5, &original_content_md5) {
if expected_md5 != actual_md5 {
return json_error_with_query(
StatusCode::BAD_REQUEST,
format!(
"Content-MD5 did not match md5 of file data expected [{}] received [{}] size {}",
expected_md5, actual_md5, original_data_size
),
Some(&query),
);
}
if let (Some(expected_md5), Some(actual_md5)) = (&content_md5, &original_content_md5)
&& expected_md5 != actual_md5
{
return json_error_with_query(
StatusCode::BAD_REQUEST,
format!(
"Content-MD5 did not match md5 of file data expected [{}] received [{}] size {}",
expected_md5, actual_md5, original_data_size
),
Some(&query),
);
}
let now = std::time::SystemTime::now()
@@ -2577,7 +2556,7 @@ pub async fn post_handler(
cookie,
data_size: final_data.len() as u32,
data: final_data,
last_modified: last_modified,
last_modified,
..Needle::default()
};
n.set_has_last_modified_date();
@@ -2595,22 +2574,21 @@ pub async fn post_handler(
}
// Set TTL on needle
if let Some(ref t) = ttl {
if !t.is_empty() {
n.ttl = Some(*t);
n.set_has_ttl();
}
if let Some(ref t) = ttl
&& !t.is_empty()
{
n.ttl = Some(*t);
n.set_has_ttl();
}
// Set pairs on needle
if !pair_map.is_empty() {
if let Ok(pairs_json) = serde_json::to_vec(&pair_map) {
if pairs_json.len() < 65536 {
n.pairs_size = pairs_json.len() as u16;
n.pairs = pairs_json;
n.set_has_pairs();
}
}
if !pair_map.is_empty()
&& let Ok(pairs_json) = serde_json::to_vec(&pair_map)
&& pairs_json.len() < 65536
{
n.pairs_size = pairs_json.len() as u16;
n.pairs = pairs_json;
n.set_has_pairs();
}
// Set filename on needle (matches Go: if len(pu.FileName) < 256)
@@ -2639,9 +2617,9 @@ pub async fn post_handler(
if !is_replicate && write_result.is_ok() && !state.master_url.is_empty() {
let needs_replication = {
let store = state.store.read().unwrap();
store.find_volume(vid).map_or(false, |(_, v)| {
v.super_block.replica_placement.get_copy_count() > 1
})
store
.find_volume(vid)
.is_some_and(|(_, v)| v.super_block.replica_placement.get_copy_count() > 1)
};
if needs_replication {
let state_clone = state.clone();
@@ -2664,7 +2642,7 @@ pub async fn post_handler(
let replication_result = replication
.await
.map_err(|e| format!("replication task failed: {}", e))
.and_then(|result| result);
.flatten();
if let Err(e) = replication_result {
tracing::error!("replicated write failed: {}", e);
return json_error_with_query(
@@ -2758,11 +2736,12 @@ pub async fn delete_handler(
// JWT check for writes (deletes use write key)
let file_id = extract_file_id(&path);
let token = extract_jwt(&headers, request.uri());
if let Err(_) = state
if state
.guard
.read()
.unwrap()
.check_jwt_for_file(token.as_deref(), &file_id, true)
.is_err()
{
return json_error_with_query(StatusCode::UNAUTHORIZED, "wrong jwt", Some(&del_query));
}
@@ -2793,14 +2772,14 @@ pub async fn delete_handler(
let count = ec_needle.data_size as i64;
// Step 3: Journal the delete
let mut store = state.store.write().unwrap();
if let Some(ecv) = store.find_ec_volume_mut(vid) {
if let Err(e) = ecv.journal_delete(needle_id) {
return json_error_with_query(
StatusCode::INTERNAL_SERVER_ERROR,
format!("Deletion Failed: {}", e),
Some(&del_query),
);
}
if let Some(ecv) = store.find_ec_volume_mut(vid)
&& let Err(e) = ecv.journal_delete(needle_id)
{
return json_error_with_query(
StatusCode::INTERNAL_SERVER_ERROR,
format!("Deletion Failed: {}", e),
Some(&del_query),
);
}
let result = DeleteResult { size: count };
return json_response_with_params(
@@ -2946,12 +2925,12 @@ pub async fn delete_handler(
if !is_replicate && delete_result.is_ok() && !state.master_url.is_empty() {
let needs_replication = {
let store = state.store.read().unwrap();
store.find_volume(vid).map_or(false, |(_, v)| {
v.super_block.replica_placement.get_copy_count() > 1
})
store
.find_volume(vid)
.is_some_and(|(_, v)| v.super_block.replica_placement.get_copy_count() > 1)
};
if needs_replication {
if let Err(e) = do_replicated_request(
if needs_replication
&& let Err(e) = do_replicated_request(
&state,
vid.0,
Method::DELETE,
@@ -2961,14 +2940,13 @@ pub async fn delete_handler(
None,
)
.await
{
tracing::error!("replicated delete failed: {}", e);
return json_error_with_query(
StatusCode::INTERNAL_SERVER_ERROR,
format!("replication failed: {}", e),
Some(&del_query),
);
}
{
tracing::error!("replicated delete failed: {}", e);
return json_error_with_query(
StatusCode::INTERNAL_SERVER_ERROR,
format!("replication failed: {}", e),
Some(&del_query),
);
}
}
@@ -3234,7 +3212,6 @@ pub async fn ui_handler(State(state): State<Arc<VolumeServerState>>) -> Response
// ============================================================================
#[derive(Deserialize)]
#[allow(dead_code)]
struct ChunkManifest {
#[serde(default)]
name: String,
@@ -3254,6 +3231,7 @@ struct ChunkInfo {
}
/// Try to expand a chunk manifest needle. Returns None if manifest can't be parsed.
#[expect(clippy::too_many_arguments)]
async fn try_expand_chunk_manifest(
state: &Arc<VolumeServerState>,
n: &Needle,
@@ -3382,53 +3360,53 @@ async fn try_expand_chunk_manifest(
response_headers.insert(header::ACCEPT_RANGES, "bytes".parse().unwrap());
// Last-Modified — Go sets this on the response writer before tryHandleChunkedFile
if let Some(lm) = last_modified_str {
if let Ok(hval) = lm.parse() {
response_headers.insert(header::LAST_MODIFIED, hval);
}
if let Some(lm) = last_modified_str
&& let Ok(hval) = lm.parse()
{
response_headers.insert(header::LAST_MODIFIED, hval);
}
// Pairs — Go sets needle pairs on the response writer before tryHandleChunkedFile
if n.has_pairs() && !n.pairs.is_empty() {
if let Ok(pair_map) =
if n.has_pairs()
&& !n.pairs.is_empty()
&& let Ok(pair_map) =
serde_json::from_slice::<std::collections::HashMap<String, String>>(&n.pairs)
{
for (k, v) in &pair_map {
if let (Ok(hname), Ok(hval)) = (
axum::http::HeaderName::from_bytes(k.as_bytes()),
axum::http::HeaderValue::from_str(v),
) {
response_headers.insert(hname, hval);
}
{
for (k, v) in &pair_map {
if let (Ok(hname), Ok(hval)) = (
axum::http::HeaderName::from_bytes(k.as_bytes()),
axum::http::HeaderValue::from_str(v),
) {
response_headers.insert(hname, hval);
}
}
}
// S3 response passthrough headers — Go sets these via AdjustPassthroughHeaders
if let Some(ref cc) = query.response_cache_control {
if let Ok(hval) = cc.parse() {
response_headers.insert(header::CACHE_CONTROL, hval);
}
if let Some(ref cc) = query.response_cache_control
&& let Ok(hval) = cc.parse()
{
response_headers.insert(header::CACHE_CONTROL, hval);
}
if let Some(ref ce) = query.response_content_encoding {
if let Ok(hval) = ce.parse() {
response_headers.insert(header::CONTENT_ENCODING, hval);
}
if let Some(ref ce) = query.response_content_encoding
&& let Ok(hval) = ce.parse()
{
response_headers.insert(header::CONTENT_ENCODING, hval);
}
if let Some(ref exp) = query.response_expires {
if let Ok(hval) = exp.parse() {
response_headers.insert(header::EXPIRES, hval);
}
if let Some(ref exp) = query.response_expires
&& let Ok(hval) = exp.parse()
{
response_headers.insert(header::EXPIRES, hval);
}
if let Some(ref cl) = query.response_content_language {
if let Ok(hval) = cl.parse() {
response_headers.insert("Content-Language", hval);
}
if let Some(ref cl) = query.response_content_language
&& let Ok(hval) = cl.parse()
{
response_headers.insert("Content-Language", hval);
}
if let Some(ref cd) = query.response_content_disposition {
if let Ok(hval) = cd.parse() {
response_headers.insert(header::CONTENT_DISPOSITION, hval);
}
if let Some(ref cd) = query.response_content_disposition
&& let Ok(hval) = cd.parse()
{
response_headers.insert(header::CONTENT_DISPOSITION, hval);
}
// Content-Disposition
@@ -3459,7 +3437,6 @@ async fn try_expand_chunk_manifest(
} else {
String::new()
};
let mut result = result;
if is_image_crop_ext(&cm_ext) {
result = maybe_crop_image(&result, &cm_ext, query);
}
@@ -3735,33 +3712,33 @@ fn extract_jwt(headers: &HeaderMap, uri: &axum::http::Uri) -> Option<String> {
// 1. Check ?jwt= query parameter
if let Some(query) = uri.query() {
for pair in query.split('&') {
if let Some(value) = pair.strip_prefix("jwt=") {
if !value.is_empty() {
return Some(value.to_string());
}
if let Some(value) = pair.strip_prefix("jwt=")
&& !value.is_empty()
{
return Some(value.to_string());
}
}
}
// 2. Check Authorization: Bearer <token> (case-insensitive prefix)
if let Some(auth) = headers.get(header::AUTHORIZATION) {
if let Ok(auth_str) = auth.to_str() {
if auth_str.len() > 7 && auth_str[..7].eq_ignore_ascii_case("bearer ") {
return Some(auth_str[7..].to_string());
}
}
if let Some(auth) = headers.get(header::AUTHORIZATION)
&& let Ok(auth_str) = auth.to_str()
&& auth_str.len() > 7
&& auth_str[..7].eq_ignore_ascii_case("bearer ")
{
return Some(auth_str[7..].to_string());
}
// 3. Check Cookie
if let Some(cookie_header) = headers.get(header::COOKIE) {
if let Ok(cookie_str) = cookie_header.to_str() {
for cookie in cookie_str.split(';') {
let cookie = cookie.trim();
if let Some(value) = cookie.strip_prefix("AT=") {
if !value.is_empty() {
return Some(value.to_string());
}
}
if let Some(cookie_header) = headers.get(header::COOKIE)
&& let Ok(cookie_str) = cookie_header.to_str()
{
for cookie in cookie_str.split(';') {
let cookie = cookie.trim();
if let Some(value) = cookie.strip_prefix("AT=")
&& !value.is_empty()
{
return Some(value.to_string());
}
}
}
+7 -8
View File
@@ -189,10 +189,10 @@ pub async fn run_heartbeat_with_state(
pub fn to_grpc_address(master_addr: &str) -> String {
if let Some((host, port_str)) = master_addr.rsplit_once(':') {
// "host:port.grpcPort" — the part after the last '.' is the gRPC port.
if let Some((_, grpc_port)) = port_str.rsplit_once('.') {
if grpc_port.parse::<u16>().is_ok() {
return format!("{}:{}", host, grpc_port);
}
if let Some((_, grpc_port)) = port_str.rsplit_once('.')
&& grpc_port.parse::<u16>().is_ok()
{
return format!("{}:{}", host, grpc_port);
}
if let Ok(port) = port_str.parse::<u16>() {
let grpc_port = port + 10000;
@@ -922,10 +922,9 @@ fn build_heartbeat_with_ec_status(
let mut effective_max_count = loc.max_volume_count.load(Ordering::Relaxed);
if loc.is_disk_space_low.load(Ordering::Relaxed) {
let used_slots = loc.volumes_len() as i32
+ ((loc.ec_shard_count()
+ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT
- 1)
/ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
+ loc
.ec_shard_count()
.div_ceil(crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
as i32;
effective_max_count = used_slots;
}
+97 -122
View File
@@ -236,8 +236,10 @@ pub async fn read_ec_shard_needle_distributed(
));
}
let mut n = Needle::default();
n.id = needle_id;
let mut n = Needle {
id: needle_id,
..Needle::default()
};
n.read_bytes(
&bytes,
snapshot.offset.to_actual_offset(),
@@ -309,7 +311,7 @@ pub async fn scrub_ec_volume_distributed(
// one scrub — the same race the vanished-volume policy exists to hide.
// The descriptor outlives the name, the same way the checksum plan's
// shard handles do.
let ecx_walk = fs::File::open(&ecv.ecx_file_name());
let ecx_walk = fs::File::open(ecv.ecx_file_name());
// Encode-run identity of the volume this scrub started against. The
// per-needle `scrub_snapshot_under_lock` re-resolves the volume by id
// under a fresh guard, so a teardown-and-remount of the same vid between
@@ -419,11 +421,10 @@ pub async fn scrub_ec_volume_distributed(
0,
Vec::new(),
vec![format!("EC volume id {} not found", vid.0)],
)
);
}
};
let map = ecv.shard_locations.read().unwrap().clone();
map
ecv.shard_locations.read().unwrap().clone()
};
// Walk the .ecx (private fd captured under the lock, no lock held) for the
@@ -812,9 +813,9 @@ fn needs_refresh(
let ttl = if stale || shard_count < data_shards {
Duration::from_secs(11)
} else if shard_count == total_shards {
Duration::from_secs(37 * 60)
Duration::from_mins(37)
} else {
Duration::from_secs(7 * 60)
Duration::from_mins(7)
};
age >= ttl
}
@@ -868,22 +869,19 @@ async fn cached_lookup_ec_shard_locations(
}
};
if master.is_empty() {
return Err(io::Error::new(
io::ErrorKind::Other,
"no master configured for ec shard lookup",
));
return Err(io::Error::other("no master configured for ec shard lookup"));
}
let grpc_addr =
parse_grpc_address(&master).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let endpoint = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?;
.map_err(|e| io::Error::other(e.to_string()))?;
let channel = endpoint
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(10))
.connect()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("master connect: {}", e)))?;
.map_err(|e| io::Error::other(format!("master connect: {}", e)))?;
let mut client = SeaweedClient::with_interceptor(channel, outgoing_request_id_interceptor)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
@@ -892,7 +890,7 @@ async fn cached_lookup_ec_shard_locations(
let resp = client
.lookup_ec_volume(Request::new(LookupEcVolumeRequest { volume_id: vid.0 }))
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("lookup_ec_volume: {}", e)))?;
.map_err(|e| io::Error::other(format!("lookup_ec_volume: {}", e)))?;
let resp = resp.into_inner();
let mut out = HashMap::new();
@@ -936,11 +934,11 @@ fn write_back_shard_locations(
/// Resolve the runtime matching the scrub's anchor encode generation, not the
/// first-match `find_ec_volume`. When `expected_encode_ts_ns` is 0 (legacy or
/// pre-feature), falls back to first-match so existing behavior is preserved.
fn find_ec_volume_for_scrub<'a>(
store: &'a crate::storage::store::Store,
fn find_ec_volume_for_scrub(
store: &crate::storage::store::Store,
vid: VolumeId,
expected_encode_ts_ns: i64,
) -> Option<&'a crate::storage::erasure_coding::EcVolume> {
) -> Option<&crate::storage::erasure_coding::EcVolume> {
if expected_encode_ts_ns != 0 {
store
.find_all_ec_volumes(vid)
@@ -959,16 +957,17 @@ fn format_location_as_server_address(loc: &master_pb::Location) -> String {
.url
.trim_start_matches("http://")
.trim_start_matches("https://");
if loc.grpc_port > 0 {
if let Some((host, http_port)) = raw.rsplit_once(':') {
return format!("{}:{}.{}", host, http_port, loc.grpc_port);
}
if loc.grpc_port > 0
&& let Some((host, http_port)) = raw.rsplit_once(':')
{
return format!("{}:{}.{}", host, http_port, loc.grpc_port);
}
raw.to_string()
}
/// Try direct peer read; on failure, reconstruct via Reed-Solomon
/// from the other shards. Mirrors `readOneEcShardInterval`'s tail.
#[expect(clippy::too_many_arguments)]
async fn fetch_one_interval(
state: &Arc<VolumeServerState>,
vid: VolumeId,
@@ -982,35 +981,35 @@ async fn fetch_one_interval(
expected_encode_ts_ns: i64,
) -> io::Result<(Vec<u8>, bool)> {
// Direct peer read against the cached locations for this shard.
if let Some(sources) = shard_locations.get(&shard_id) {
if !sources.is_empty() {
match read_remote_ec_shard_interval(
state,
sources,
vid,
needle_id,
shard_id,
shard_offset,
size,
expected_encode_ts_ns,
)
.await
{
// A deleted needle short-circuits: don't reconstruct (every shard
// would report deleted), let the caller return "deleted".
Ok((buf, is_deleted)) => return Ok((buf, is_deleted)),
Err(e) => {
tracing::debug!(
"direct read ec shard {}.{} from {:?} failed: {} — will reconstruct",
vid.0,
shard_id,
sources,
e
);
// Reconstruction below skips this very shard, so nothing else
// invalidates the location that just failed.
mark_shard_locations_stale(state, vid);
}
if let Some(sources) = shard_locations.get(&shard_id)
&& !sources.is_empty()
{
match read_remote_ec_shard_interval(
state,
sources,
vid,
needle_id,
shard_id,
shard_offset,
size,
expected_encode_ts_ns,
)
.await
{
// A deleted needle short-circuits: don't reconstruct (every shard
// would report deleted), let the caller return "deleted".
Ok((buf, is_deleted)) => return Ok((buf, is_deleted)),
Err(e) => {
tracing::debug!(
"direct read ec shard {}.{} from {:?} failed: {} — will reconstruct",
vid.0,
shard_id,
sources,
e
);
// Reconstruction below skips this very shard, so nothing else
// invalidates the location that just failed.
mark_shard_locations_stale(state, vid);
}
}
}
@@ -1032,6 +1031,7 @@ async fn fetch_one_interval(
.await
}
#[expect(clippy::too_many_arguments)]
async fn read_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
sources: &[String],
@@ -1068,6 +1068,7 @@ async fn read_remote_ec_shard_interval(
}))
}
#[expect(clippy::too_many_arguments)]
async fn do_read_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
source: &str,
@@ -1081,18 +1082,13 @@ async fn do_read_remote_ec_shard_interval(
let grpc_addr =
parse_grpc_address(source).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let endpoint = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?;
.map_err(|e| io::Error::other(e.to_string()))?;
let channel = endpoint
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(30))
.connect()
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("connect to {}: {}", source, e),
)
})?;
.map_err(|e| io::Error::other(format!("connect to {}: {}", source, e)))?;
// TODO(grpc-jwt): clusters with `jwt.signing.key` configured will
// reject peer-to-peer VolumeEcShardRead calls until the Rust
@@ -1118,13 +1114,10 @@ async fn do_read_remote_ec_shard_interval(
.volume_ec_shard_read(Request::new(req))
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"volume_ec_shard_read {}.{} from {}: {}",
vid.0, shard_id, source, e
),
)
io::Error::other(format!(
"volume_ec_shard_read {}.{} from {}: {}",
vid.0, shard_id, source, e
))
})?;
let mut stream = resp.into_inner();
@@ -1133,19 +1126,16 @@ async fn do_read_remote_ec_shard_interval(
while let Some(msg) = stream
.message()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("recv: {}", e)))?
.map_err(|e| io::Error::other(format!("recv: {}", e)))?
{
// Validate the served shard's identity client-side, so the guard holds even
// against a pre-upgrade server that ignored the request field (returns 0).
// A mismatch fails the read; the caller recovers from parity.
if expected_encode_ts_ns != 0 && msg.encode_ts_ns != expected_encode_ts_ns {
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"ec shard {}.{} from {} belongs to a different encode run (want {} got {})",
vid.0, shard_id, source, expected_encode_ts_ns, msg.encode_ts_ns
),
));
return Err(io::Error::other(format!(
"ec shard {}.{} from {} belongs to a different encode run (want {} got {})",
vid.0, shard_id, source, expected_encode_ts_ns, msg.encode_ts_ns
)));
}
if msg.is_deleted {
is_deleted = true;
@@ -1181,6 +1171,7 @@ async fn do_read_remote_ec_shard_interval(
Ok((out, false))
}
#[expect(clippy::too_many_arguments)]
async fn recover_one_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
vid: VolumeId,
@@ -1195,7 +1186,7 @@ async fn recover_one_remote_ec_shard_interval(
) -> io::Result<(Vec<u8>, bool)> {
let total_shards = data_shards + parity_shards;
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
// Charge the buffers this recovery is about to hold against the budget, so a
// burst of them queues here rather than on the heap. An interval whose
@@ -1205,13 +1196,10 @@ async fn recover_one_remote_ec_shard_interval(
.acquire_many((size * data_shards).min(EC_RECOVER_BUDGET) as u32)
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"ec recover budget for shard {}.{}: {}",
vid.0, shard_id_to_recover, e
),
)
io::Error::other(format!(
"ec recover budget for shard {}.{}: {}",
vid.0, shard_id_to_recover, e
))
})?;
let mut bufs: Vec<Option<Vec<u8>>> = vec![None; total_shards];
@@ -1224,7 +1212,7 @@ async fn recover_one_remote_ec_shard_interval(
let mut available = 0usize;
{
let store = state.store.read().unwrap();
for sid in 0..total_shards {
for (sid, slot) in bufs.iter_mut().enumerate() {
if available >= data_shards {
break;
}
@@ -1251,7 +1239,7 @@ async fn recover_one_remote_ec_shard_interval(
.map(|n| n == size)
.unwrap_or(false)
{
bufs[sid] = Some(buf);
*slot = Some(buf);
available += 1;
}
}
@@ -1335,34 +1323,25 @@ async fn recover_one_remote_ec_shard_interval(
if any_deleted {
return Ok((Vec::new(), true));
}
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"cannot recover ec shard {}.{}: only {} shards available, need at least {}",
vid.0, shard_id_to_recover, available, data_shards
),
));
return Err(io::Error::other(format!(
"cannot recover ec shard {}.{}: only {} shards available, need at least {}",
vid.0, shard_id_to_recover, available, data_shards
)));
}
rs.reconstruct(&mut bufs).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"reed-solomon reconstruct ec shard {}.{}: {:?}",
vid.0, shard_id_to_recover, e
),
)
io::Error::other(format!(
"reed-solomon reconstruct ec shard {}.{}: {:?}",
vid.0, shard_id_to_recover, e
))
})?;
match bufs.into_iter().nth(shard_id_to_recover as usize).flatten() {
Some(buf) => Ok((buf, any_deleted)),
None => Err(io::Error::new(
io::ErrorKind::Other,
format!(
"reconstructed buffer for shard {}.{} missing after RS reconstruct",
vid.0, shard_id_to_recover
),
)),
None => Err(io::Error::other(format!(
"reconstructed buffer for shard {}.{} missing after RS reconstruct",
vid.0, shard_id_to_recover
))),
}
}
@@ -1504,12 +1483,12 @@ async fn fetch_ec_index_from_one_peer(
let grpc_addr =
parse_grpc_address(peer).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let channel = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?
.map_err(|e| io::Error::other(e.to_string()))?
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(30))
.connect()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("connect {}: {}", peer, e)))?;
.map_err(|e| io::Error::other(format!("connect {}: {}", peer, e)))?;
let mut client = VolumeServerClient::with_interceptor(channel, outgoing_request_id_interceptor)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE);
@@ -1530,22 +1509,19 @@ async fn fetch_ec_index_from_one_peer(
let stream = client
.copy_file(copy_req(".ecx", false))
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("copy .ecx: {}", e)))?
.map_err(|e| io::Error::other(format!("copy .ecx: {}", e)))?
.into_inner();
drain_copy_stream(stream, ecx_path, false).await?;
let meta = fs::metadata(ecx_path)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("stat copied .ecx: {}", e)))?;
let meta =
fs::metadata(ecx_path).map_err(|e| io::Error::other(format!("stat copied .ecx: {}", e)))?;
if meta.is_dir() || meta.len() == 0 {
let _ = fs::remove_file(ecx_path);
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"peer {} served an unusable .ecx (size {})",
peer,
meta.len()
),
));
return Err(io::Error::other(format!(
"peer {} served an unusable .ecx (size {})",
peer,
meta.len()
)));
}
// .ecj is the source peer's deletion journal (appended); .vif carries EC
@@ -1589,15 +1565,14 @@ async fn drain_copy_stream(
} else {
fs::File::create(dest_path)
}
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("create {}: {}", dest_path, e)))?;
.map_err(|e| io::Error::other(format!("create {}: {}", dest_path, e)))?;
while let Some(chunk) = stream
.message()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("recv {}: {}", dest_path, e)))?
.map_err(|e| io::Error::other(format!("recv {}: {}", dest_path, e)))?
{
file.write_all(&chunk.file_content).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("write {}: {}", dest_path, e))
})?;
file.write_all(&chunk.file_content)
.map_err(|e| io::Error::other(format!("write {}: {}", dest_path, e)))?;
}
Ok(())
}
+6 -7
View File
@@ -423,13 +423,12 @@ fn collect_ui_data(
shard_id: shard.shard_id,
size: shard_size,
});
if created_at == "-" {
if let Ok(metadata) = std::fs::metadata(shard.file_name()) {
if let Ok(modified) = metadata.modified() {
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
if created_at == "-"
&& let Ok(metadata) = std::fs::metadata(shard.file_name())
&& let Ok(modified) = metadata.modified()
{
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
let preferred_size = ec_volume.dat_file_size.max(0) as u64;
+8 -10
View File
@@ -312,16 +312,15 @@ async fn admin_store_handler(state: State<Arc<VolumeServerState>>, request: Requ
)
}
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
@@ -358,16 +357,15 @@ async fn public_store_handler(state: State<Arc<VolumeServerState>>, request: Req
}
_ => StatusCode::OK.into_response(),
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
+34 -29
View File
@@ -131,10 +131,10 @@ impl DiskLocation {
for entry in entries {
let entry = entry?;
let name = entry.file_name().into_string().unwrap_or_default();
if let Some((collection, vid)) = parse_volume_filename(&name) {
if seen.insert((collection.clone(), vid)) {
dat_files.push((collection, vid));
}
if let Some((collection, vid)) = parse_volume_filename(&name)
&& seen.insert((collection.clone(), vid))
{
dat_files.push((collection, vid));
}
}
@@ -327,10 +327,10 @@ impl DiskLocation {
.strip_suffix(".cpc")
.or_else(|| name.strip_suffix(".cpd"))
.or_else(|| name.strip_suffix(".cpx"));
if let Some(stem) = stem {
if let Some(key) = parse_collection_volume_id(stem) {
pending.insert(key);
}
if let Some(stem) = stem
&& let Some(key) = parse_collection_volume_id(stem)
{
pending.insert(key);
}
}
}
@@ -426,11 +426,16 @@ impl DiskLocation {
if shard_count == 0 {
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size) {
if actual < expected {
warn!(volume_id = vid.0, actual, expected, "shards smaller than the .dat's full encode; reclaiming the complete .dat");
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size)
&& actual < expected
{
warn!(
volume_id = vid.0,
actual,
expected,
"shards smaller than the .dat's full encode; reclaiming the complete .dat"
);
return false;
}
true
}
@@ -510,10 +515,10 @@ impl DiskLocation {
pub(crate) fn ec_generation_ts_ns(&self, collection: &str, vid: VolumeId) -> Option<i64> {
for dir in [&self.directory, &self.idx_directory] {
let vif = format!("{}.vif", volume_file_name(dir, collection, vid));
if let Ok(s) = fs::read_to_string(&vif) {
if let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s) {
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if let Ok(s) = fs::read_to_string(&vif)
&& let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s)
{
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if self.directory == self.idx_directory {
break;
@@ -542,6 +547,7 @@ impl DiskLocation {
}
/// Create a new volume in this location.
#[expect(clippy::too_many_arguments)]
pub fn create_volume(
&mut self,
vid: VolumeId,
@@ -777,18 +783,18 @@ impl DiskLocation {
pub fn has_ecx_file_on_disk(&self, collection: &str, vid: VolumeId) -> bool {
let idx_base = volume_file_name(&self.idx_directory, collection, vid);
let idx_path = format!("{}.ecx", idx_base);
if let Ok(meta) = fs::metadata(&idx_path) {
if !meta.is_dir() {
return true;
}
if let Ok(meta) = fs::metadata(&idx_path)
&& !meta.is_dir()
{
return true;
}
if self.idx_directory != self.directory {
let data_base = volume_file_name(&self.directory, collection, vid);
let data_path = format!("{}.ecx", data_base);
if let Ok(meta) = fs::metadata(&data_path) {
if !meta.is_dir() {
return true;
}
if let Ok(meta) = fs::metadata(&data_path)
&& !meta.is_dir()
{
return true;
}
}
false
@@ -1107,7 +1113,7 @@ impl DiskLocation {
/// Close all volumes.
pub fn close(&mut self) {
for (_, v) in self.volumes.iter_mut() {
for v in self.volumes.values_mut() {
v.close();
}
self.volumes.clear();
@@ -1184,10 +1190,9 @@ fn ec_data_shards_from_vif(directory: &str, idx_directory: &str, collection: &st
.and_then(|s| serde_json::from_str::<VifVolumeInfo>(&s).ok())
.and_then(|vi| vi.ec_shard_config)
.map(|c| c.data_shards as usize)
&& ds > 0
{
if ds > 0 {
return ds;
}
return ds;
}
if directory == idx_directory {
break;
@@ -164,16 +164,16 @@ pub fn remove_bitrot_sidecars(base: &str) -> io::Result<()> {
};
let mut first_err: Option<io::Error> = None;
let mut record = |res: io::Result<()>| {
if let Err(e) = res {
if first_err.is_none() {
first_err = Some(e);
}
if let Err(e) = res
&& first_err.is_none()
{
first_err = Some(e);
}
};
record(rm(format!("{}{}", base, BITROT_SIDECAR_EXT).into()));
let path = Path::new(base);
if let (Some(parent), Some(fname)) = (path.parent(), path.file_name()) {
let prefix = format!("{}{}.v", fname.to_string_lossy(), BITROT_SIDECAR_EXT);
let prefix = format!("{}{}.v", fname.display(), BITROT_SIDECAR_EXT);
match fs::read_dir(parent) {
Ok(entries) => {
for entry in entries.flatten() {
@@ -203,7 +203,7 @@ pub fn new_encode_uuid() -> Vec<u8> {
/// Reports whether `block_size` is a power of two in [1 MiB, MAX_BITROT_BLOCK_SIZE].
pub fn is_pow2_multiple_of_1mib(block_size: u32) -> bool {
block_size >= (1 << 20) && block_size <= MAX_BITROT_BLOCK_SIZE && block_size.count_ones() == 1
((1 << 20)..=MAX_BITROT_BLOCK_SIZE).contains(&block_size) && block_size.count_ones() == 1
}
/// Returns ceil(covered_size / block_size).
@@ -402,7 +402,7 @@ pub fn validate_manifest(
total
));
}
let mut seen = vec![false; MAX_SHARD_COUNT];
let mut seen = [false; MAX_SHARD_COUNT];
for s in &prot.shards {
if s.shard_id >= total as u32 {
return Err(format!(
@@ -73,6 +73,7 @@ pub fn find_dat_file_size_with_dirs(
/// must live in `dir`. For the cross-disk reconciled layout where
/// shards are split across multiple data dirs of the same node, use
/// [`write_dat_file_from_shards_with_dirs`] instead.
#[expect(clippy::too_many_arguments)]
pub fn write_dat_file_from_shards(
dir: &str,
collection: &str,
@@ -120,7 +121,7 @@ pub fn write_dat_file_from_shards(
/// size. `large_block_size`/`small_block_size` are the volume's shard
/// block layout, e.g. `EcVolume::large_block_size()` /
/// `small_block_size()` from its .vif EC config.
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
pub fn write_dat_file_from_shards_with_dirs(
dat_dir: &str,
collection: &str,
@@ -145,7 +146,7 @@ pub fn write_dat_file_from_shards_with_dirs(
)
}
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
fn write_dat_file(
dat_dir: &str,
collection: &str,
@@ -233,10 +234,10 @@ fn write_dat_file(
// Read large blocks
while encoded_remaining >= large_row_size && remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = large_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -255,10 +256,10 @@ fn write_dat_file(
// Read small blocks
while remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = small_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -324,10 +325,7 @@ pub fn write_idx_file_from_ec_index(
// and treat only NotFound as "no journal": Path::exists would also
// swallow a permission/IO error and silently skip deletions, which
// would resurrect deleted needles as live.
let mut idx_file = std::fs::OpenOptions::new()
.write(true)
.append(true)
.open(&tmp_path)?;
let mut idx_file = std::fs::OpenOptions::new().append(true).open(&tmp_path)?;
match std::fs::read(&ecj_path) {
Ok(ecj_data) => {
let count = ecj_data.len() / NEEDLE_ID_SIZE;
@@ -50,7 +50,7 @@ pub fn write_ec_files(
let dat_size = dat_file.metadata()?.len() as i64;
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
// Create shard files
let total_shards = data_shards + parity_shards;
@@ -162,7 +162,7 @@ pub fn rebuild_ec_files(
}
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
@@ -175,7 +175,7 @@ pub fn rebuild_ec_files(
let mut shard_size = 0;
for (i, shard) in shards.iter_mut().enumerate() {
if !missing_shard_ids.contains(&(i as u32)) {
if let Ok(_) = shard.open() {
if shard.open().is_ok() {
let size = shard.file_size();
if size > shard_size {
shard_size = size;
@@ -185,7 +185,7 @@ pub fn rebuild_ec_files(
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
if let Ok(_) = alt.open() {
if alt.open().is_ok() {
let size = alt.file_size();
if size > shard_size {
shard_size = size;
@@ -251,12 +251,8 @@ pub fn rebuild_ec_files(
}
// Reconstruct missing shards
rs.reconstruct(&mut buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon reconstruct: {:?}", e),
)
})?;
rs.reconstruct(&mut buffers)
.map_err(|e| io::Error::other(format!("reed-solomon reconstruct: {:?}", e)))?;
// Write recovered data into the missing shards
for i in missing_shard_ids {
@@ -296,7 +292,7 @@ pub fn verify_ec_shards(
parity_shards: usize,
) -> io::Result<(Vec<u32>, Vec<String>)> {
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<Option<EcVolumeShard>> = (0..total_shards)
@@ -378,27 +374,27 @@ pub fn verify_ec_shards(
if !read_failed {
// Need to convert Vec<Vec<u8>> to &[&[u8]] for rs.verify
let slice_ptrs: Vec<&[u8]> = buffers.iter().map(|v| v.as_slice()).collect();
if let Ok(is_valid) = rs.verify(&slice_ptrs) {
if !is_valid {
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
if let Ok(is_valid) = rs.verify(&slice_ptrs)
&& !is_valid
{
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for i in data_shards..total_shards {
verify_buffers[i].fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for buf in &mut verify_buffers[data_shards..total_shards] {
buf.fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
}
}
@@ -490,7 +486,7 @@ pub fn rebuild_ecx_file(
.collect();
for (i, shard) in shards.iter_mut().enumerate() {
if let Err(_) = shard.open() {
if shard.open().is_err() {
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
@@ -507,7 +503,7 @@ pub fn rebuild_ecx_file(
}
return Err(io::Error::new(
io::ErrorKind::NotFound,
format!("cannot open data shard for ecx rebuild"),
"cannot open data shard for ecx rebuild".to_string(),
));
}
}
@@ -515,7 +511,7 @@ pub fn rebuild_ecx_file(
// Determine total logical data size from shard sizes
let shard_size = shards.iter().map(|s| s.file_size()).max().unwrap_or(0);
let total_data_size = shard_size as i64 * data_shards as i64;
let total_data_size = shard_size * data_shards as i64;
// The volume's shard block layout: the .vif-recorded uniform block size,
// or the legacy two-tier sizes when 0. The row count comes from the shard
// length; -1 disambiguates a legacy shard that is an exact large-block
@@ -538,7 +534,7 @@ pub fn rebuild_ecx_file(
let locate_shard_size = if dat_file_size > 0 {
dat_file_size / data_shards as i64
} else {
(shard_size as i64 - 1).max(0)
(shard_size - 1).max(0)
};
// Read version from superblock (first byte of logical data)
@@ -640,7 +636,6 @@ pub fn rebuild_ecx_file(
/// Read bytes from EC data shards at a logical offset in the .dat file,
/// resolving the shard/offset through the volume's block layout via
/// locate_data — the same mapping the read path uses.
#[allow(clippy::too_many_arguments)]
fn read_from_data_shards(
shards: &[EcVolumeShard],
buf: &mut [u8],
@@ -714,7 +709,7 @@ const ENCODE_BUFFER_SIZE: usize = 256 * 1024;
/// 2. Process remaining data with small blocks
///
/// `buffer_size` must divide both block sizes.
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
pub(crate) fn encode_dat_file(
dat_file: &File,
dat_size: i64,
@@ -778,7 +773,7 @@ pub(crate) fn encode_dat_file(
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches so
/// arbitrarily large blocks never require block-sized allocations. Mirrors
/// Go's encodeData.
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
fn encode_data(
dat_file: &File,
row_offset: u64,
@@ -790,7 +785,7 @@ fn encode_data(
data_shards: usize,
) -> io::Result<()> {
let buffer_size = buffers[0].len();
if block_size % buffer_size != 0 {
if !block_size.is_multiple_of(buffer_size) {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
@@ -817,7 +812,7 @@ fn encode_data(
/// Encode one sub-batch: the same buffer-sized slice of every shard's block in
/// this row. Mirrors Go's encodeDataOneBatch.
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
fn encode_one_batch(
dat_file: &File,
offset: u64,
@@ -830,21 +825,15 @@ fn encode_one_batch(
) -> io::Result<()> {
// Read data shards from the .dat file, zero-filling past EOF — the buffers
// are reused across batches, so the tail must be cleared explicitly.
for i in 0..data_shards {
for (i, buf) in buffers[..data_shards].iter_mut().enumerate() {
let read_offset = offset + (i * block_size) as u64;
let n = read_at_most(dat_file, &mut buffers[i], read_offset)?;
for b in buffers[i][n..].iter_mut() {
*b = 0;
}
let n = read_at_most(dat_file, buf, read_offset)?;
buf[n..].fill(0);
}
// Encode parity shards
rs.encode(&mut *buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon encode: {:?}", e),
)
})?;
rs.encode(&mut *buffers)
.map_err(|e| io::Error::other(format!("reed-solomon encode: {:?}", e)))?;
// Write all shard buffers to files and feed the same bytes to each
// shard's bitrot checksum builder, keeping covered_size == on-disk length.
@@ -78,7 +78,7 @@ impl EcVolumeShard {
let file = self
.ecd_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
#[cfg(unix)]
{
@@ -102,7 +102,7 @@ impl EcVolumeShard {
let file = self
.ecd_file
.as_mut()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
file.write_all(data)?;
self.ecd_file_size += data.len() as i64;
Ok(())
@@ -123,7 +123,7 @@ impl EcVolumeShard {
pub fn try_clone_file(&self) -> io::Result<File> {
self.ecd_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?
.ok_or_else(|| io::Error::other("shard file not open"))?
.try_clone()
}
@@ -565,32 +565,31 @@ impl EcVolume {
// A sidecar written for THIS generation that contradicts the volume's
// geometry is not "no protection" — it says the layout the volume is
// about to serve reads with is wrong. Fail the mount.
if let Ok(prot) = &loaded {
if prot.generation == generation
&& !ec_bitrot::geometry_matches(
prot,
self.data_shards as usize,
self.parity_shards as usize,
if let Ok(prot) = &loaded
&& prot.generation == generation
&& !ec_bitrot::geometry_matches(
prot,
self.data_shards as usize,
self.parity_shards as usize,
self.block_size,
)
{
let cfg = prot.ec_shard_config.as_ref();
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"ec volume {} generation {}: {} records layout {}+{} block {} but the volume is mounted as {}+{} block {}; refusing to serve one of the two layouts",
self.volume_id.0,
generation,
path,
cfg.map(|c| c.data_shards).unwrap_or(0),
cfg.map(|c| c.parity_shards).unwrap_or(0),
cfg.map(|c| c.block_size).unwrap_or(0),
self.data_shards,
self.parity_shards,
self.block_size,
)
{
let cfg = prot.ec_shard_config.as_ref();
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"ec volume {} generation {}: {} records layout {}+{} block {} but the volume is mounted as {}+{} block {}; refusing to serve one of the two layouts",
self.volume_id.0,
generation,
path,
cfg.map(|c| c.data_shards).unwrap_or(0),
cfg.map(|c| c.parity_shards).unwrap_or(0),
cfg.map(|c| c.block_size).unwrap_or(0),
self.data_shards,
self.parity_shards,
self.block_size,
),
));
}
),
));
}
let status = ec_bitrot::resolve_status(
&loaded,
@@ -698,7 +697,7 @@ impl EcVolume {
let mut set = self
.deleted_needles
.write()
.map_err(|_| io::Error::new(io::ErrorKind::Other, "deleted_needles lock poisoned"))?;
.map_err(|_| io::Error::other("deleted_needles lock poisoned"))?;
let mut off: i64 = 0;
while off + NEEDLE_ID_SIZE as i64 <= self.ecj_file_size {
#[cfg(unix)]
@@ -744,7 +743,6 @@ impl EcVolume {
// ---- File names ----
#[allow(dead_code)]
fn base_name(&self) -> String {
crate::storage::volume::volume_file_name(&self.dir, &self.collection, self.volume_id)
}
@@ -826,10 +824,8 @@ impl EcVolume {
/// default to the physical location's disk type.
pub fn set_disk_type(&mut self, d: DiskType) {
self.disk_type = d.clone();
for slot in self.shards.iter_mut() {
if let Some(shard) = slot {
shard.disk_type = d.clone();
}
for shard in self.shards.iter_mut().flatten() {
shard.disk_type = d.clone();
}
}
@@ -986,21 +982,21 @@ impl EcVolume {
pub fn check_read_write_error(&self, err: Option<&io::Error>) {
use std::sync::atomic::Ordering;
if let Some(e) = err {
if crate::storage::volume::is_storage_io_error(e) {
self.io_error_count.fetch_add(1, Ordering::Relaxed);
if let Ok(mut guard) = self.last_io_error.lock() {
*guard = Some(e.to_string());
}
crate::metrics::STORAGE_IO_ERROR_COUNTER.inc();
return;
if let Some(e) = err
&& crate::storage::volume::is_storage_io_error(e)
{
self.io_error_count.fetch_add(1, Ordering::Relaxed);
if let Ok(mut guard) = self.last_io_error.lock() {
*guard = Some(e.to_string());
}
crate::metrics::STORAGE_IO_ERROR_COUNTER.inc();
return;
}
self.io_error_count.store(0, Ordering::Relaxed);
if let Ok(mut guard) = self.last_io_error.lock() {
if guard.is_some() {
*guard = None;
}
if let Ok(mut guard) = self.last_io_error.lock()
&& guard.is_some()
{
*guard = None;
}
}
@@ -1033,7 +1029,7 @@ impl EcVolume {
let ecx_file = self
.ecx_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "ecx file not open"))?;
.ok_or_else(|| io::Error::other("ecx file not open"))?;
let entry_count = self.ecx_file_size as usize / NEEDLE_MAP_ENTRY_SIZE;
if entry_count == 0 {
@@ -1253,10 +1249,8 @@ impl EcVolume {
/// Get the size of a single shard (all shards are the same size).
fn shard_file_size(&self) -> i64 {
for shard in &self.shards {
if let Some(s) = shard {
return s.file_size();
}
if let Some(s) = self.shards.iter().flatten().next() {
return s.file_size();
}
0
}
@@ -1363,13 +1357,10 @@ impl EcVolume {
/// the index (ignored by callers) and an error on IO failure.
fn tombstone_ecx_entry(&self, needle_id: NeedleId) -> io::Result<bool> {
let ecx_file = self.ecx_file.as_ref().ok_or_else(|| {
io::Error::new(
io::ErrorKind::Other,
format!(
"ec volume {} has no open .ecx file (closed or corrupt)",
self.volume_id.0
),
)
io::Error::other(format!(
"ec volume {} has no open .ecx file (closed or corrupt)",
self.volume_id.0
))
})?;
let entry_count = self.ecx_file_size as usize / NEEDLE_MAP_ENTRY_SIZE;
@@ -1416,7 +1407,7 @@ impl EcVolume {
/// `deleted_needles` instead). The rebuild is atomic with respect to
/// the journal: if any individual write fails the .ecj file is left
/// in place and the error is propagated so tombstones are not lost.
#[allow(dead_code)]
#[expect(dead_code, reason = "no caller yet; see the doc comment")]
fn rebuild_ecx_from_journal(&mut self) -> io::Result<()> {
let ecj_path = self.ecj_file_name();
if !std::path::Path::new(&ecj_path).exists() {
@@ -1510,7 +1501,7 @@ impl EcVolume {
let ecj_file = self
.ecj_file
.as_mut()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "ecj file not open"))?;
.ok_or_else(|| io::Error::other("ecj file not open"))?;
let mut buf = [0u8; NEEDLE_ID_SIZE];
needle_id.to_bytes(&mut buf);
ecj_file.write_all(&buf).and_then(|_| ecj_file.sync_all())
@@ -1528,15 +1519,15 @@ impl EcVolume {
// write_all may have extended the file on disk before
// sync_all failed; truncate back to the known-good size so
// the on-disk journal never drifts past `deleted_needles`.
if let Some(ecj) = self.ecj_file.as_mut() {
if let Err(trunc_err) = ecj.set_len(prev_ecj_size as u64) {
tracing::error!(
volume_id = self.volume_id.0,
needle_id = needle_id.0,
truncate_error = %trunc_err,
"failed to truncate ecj after append failure"
);
}
if let Some(ecj) = self.ecj_file.as_mut()
&& let Err(trunc_err) = ecj.set_len(prev_ecj_size as u64)
{
tracing::error!(
volume_id = self.volume_id.0,
needle_id = needle_id.0,
truncate_error = %trunc_err,
"failed to truncate ecj after append failure"
);
}
Err(e)
}
@@ -1550,7 +1541,7 @@ impl EcVolume {
let ecx_file = self
.ecx_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "ecx file not open"))?;
.ok_or_else(|| io::Error::other("ecx file not open"))?;
let entry_count = self.ecx_file_size as usize / NEEDLE_MAP_ENTRY_SIZE;
if entry_count == 0 {
return Ok(None);
@@ -1590,31 +1581,32 @@ impl EcVolume {
if cookie.0 != 0 {
// Try to read the needle's cookie from the EC shards to validate
// Look up the needle in ecx index to find its offset, then read header from shard
if let Ok(Some((offset, size))) = self.find_needle_from_ecx(needle_id) {
if !size.is_deleted() && !offset.is_zero() {
let actual_offset = offset.to_actual_offset() as u64;
// Determine which shard contains this offset and read the cookie
let shard_size = self
.shards
.iter()
.filter_map(|s| s.as_ref())
.map(|s| s.file_size())
.next()
.unwrap_or(0) as u64;
if shard_size > 0 {
let shard_id = (actual_offset / shard_size) as usize;
let shard_offset = actual_offset % shard_size;
if let Some(Some(shard)) = self.shards.get(shard_id) {
let mut header_buf = [0u8; 4]; // cookie is first 4 bytes of needle
if shard.read_at(&mut header_buf, shard_offset).is_ok() {
let needle_cookie =
crate::storage::types::Cookie(u32::from_be_bytes(header_buf));
if needle_cookie != cookie {
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!("unexpected cookie {:x}", cookie.0),
));
}
if let Ok(Some((offset, size))) = self.find_needle_from_ecx(needle_id)
&& !size.is_deleted()
&& !offset.is_zero()
{
let actual_offset = offset.to_actual_offset() as u64;
// Determine which shard contains this offset and read the cookie
let shard_size = self
.shards
.iter()
.filter_map(|s| s.as_ref())
.map(|s| s.file_size())
.next()
.unwrap_or(0) as u64;
if let Some(shard_id) = actual_offset.checked_div(shard_size) {
let shard_id = shard_id as usize;
let shard_offset = actual_offset % shard_size;
if let Some(Some(shard)) = self.shards.get(shard_id) {
let mut header_buf = [0u8; 4]; // cookie is first 4 bytes of needle
if shard.read_at(&mut header_buf, shard_offset).is_ok() {
let needle_cookie =
crate::storage::types::Cookie(u32::from_be_bytes(header_buf));
if needle_cookie != cookie {
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!("unexpected cookie {:x}", cookie.0),
));
}
}
}
@@ -3742,10 +3734,10 @@ pub(crate) fn merge_ec_runtimes<'a>(runtimes: &[&'a EcVolume]) -> Option<MergedE
let mut slots: Vec<Option<(&'a EcVolume, &'a EcVolumeShard)>> = vec![None; width];
for v in &merged {
for (id, slot) in v.shards.iter().enumerate() {
if let Some(shard) = slot.as_ref() {
if slots[id].is_none() {
slots[id] = Some((*v, shard));
}
if let Some(shard) = slot.as_ref()
&& slots[id].is_none()
{
slots[id] = Some((*v, shard));
}
}
}
@@ -4366,13 +4358,10 @@ impl EcLocalScrubPlan {
if read != want {
// Like Go, returning from the walk callback aborts the scan.
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"expected {} bytes for needle {} on volume {}, got {}",
want, id.0, volume_id.0, read
),
));
return Err(io::Error::other(format!(
"expected {} bytes for needle {} on volume {}, got {}",
want, id.0, volume_id.0, read
)));
}
// Only a fully-local needle can be reassembled and CRC-checked.
@@ -4405,7 +4394,7 @@ impl EcLocalScrubPlan {
.filter_map(|sid| shards.get(*sid as usize).and_then(|s| s.as_ref()))
.map(|s| s.info.clone())
.collect();
broken.sort_by(|a, b| a.shard_id.cmp(&b.shard_id));
broken.sort_by_key(|a| a.shard_id);
(count, broken, errs)
}
+4 -2
View File
@@ -21,7 +21,7 @@ impl CRC {
/// Legacy `.Value()` function — deprecated in Go but needed for backward compat check.
/// Formula: (crc >> 15 | crc << 17) + 0xa282ead8
pub fn legacy_value(&self) -> u32 {
(self.0 >> 15 | self.0 << 17).wrapping_add(0xa282ead8)
self.0.rotate_right(15).wrapping_add(0xa282ead8)
}
}
@@ -67,7 +67,9 @@ mod tests {
fn test_crc_legacy_value() {
let crc = CRC(0x12345678);
let v = crc.legacy_value();
let expected = (0x12345678u32 >> 15 | 0x12345678u32 << 17).wrapping_add(0xa282ead8);
// (0x12345678 >> 15 | 0x12345678 << 17) + 0xa282ead8, worked out by hand so
// the test checks the rotate rather than restating it.
let expected = 0x4f730f40_u32;
assert_eq!(v, expected);
}
}
+1
View File
@@ -1,4 +1,5 @@
pub mod crc;
#[expect(clippy::module_inception, reason = "needle/needle.rs mirrors the Go package layout")]
pub mod needle;
pub mod ttl;
+21 -15
View File
@@ -560,7 +560,7 @@ impl Needle {
// Padding to 8-byte alignment
let padding = padding_length(self.size, version).0 as usize;
buf.extend(std::iter::repeat(0u8).take(padding));
buf.extend(std::iter::repeat_n(0u8, padding));
buf
}
@@ -824,11 +824,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v3() {
let mut n = Needle::default();
n.cookie = Cookie(42);
n.id = NeedleId(100);
n.data = b"hello world".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(42),
id: NeedleId(100),
data: b"hello world".to_vec(),
flags: 0,
..Needle::default()
};
n.set_has_name();
n.name = b"test.txt".to_vec();
n.name_size = 8;
@@ -867,11 +869,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v2() {
let mut n = Needle::default();
n.cookie = Cookie(77);
n.id = NeedleId(200);
n.data = b"data v2".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(77),
id: NeedleId(200),
data: b"data v2".to_vec(),
flags: 0,
..Needle::default()
};
let bytes = n.write_bytes(VERSION_2);
let expected_size = get_actual_size(n.size, VERSION_2);
@@ -886,10 +890,12 @@ mod tests {
#[test]
fn test_read_bytes_meta_only_handles_tombstone_v3() {
let mut tombstone = Needle::default();
tombstone.cookie = Cookie(0x1234abcd);
tombstone.id = NeedleId(300);
tombstone.append_at_ns = 999_999;
let mut tombstone = Needle {
cookie: Cookie(0x1234abcd),
id: NeedleId(300),
append_at_ns: 999_999,
..Needle::default()
};
let bytes = tombstone.write_bytes(VERSION_3);
+50 -17
View File
@@ -81,7 +81,7 @@ impl TTL {
return Ok(TTL::EMPTY);
}
let last_byte = s.as_bytes()[s.len() - 1];
let (num_str, unit_byte) = if last_byte >= b'0' && last_byte <= b'9' {
let (num_str, unit_byte) = if last_byte.is_ascii_digit() {
// All digits — default to minutes (matching Go)
(s, b'm')
} else {
@@ -144,40 +144,73 @@ fn fit_ttl_count(count: u32, unit: u8) -> TTL {
const MINUTE_SECS: u64 = 60;
// First pass: try exact fits from largest to smallest
if seconds % YEAR_SECS == 0 && seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
if seconds.is_multiple_of(YEAR_SECS) && seconds / YEAR_SECS < 256 {
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
if seconds % MONTH_SECS == 0 && seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
if seconds.is_multiple_of(MONTH_SECS) && seconds / MONTH_SECS < 256 {
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds % WEEK_SECS == 0 && seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
if seconds.is_multiple_of(WEEK_SECS) && seconds / WEEK_SECS < 256 {
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds % DAY_SECS == 0 && seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
if seconds.is_multiple_of(DAY_SECS) && seconds / DAY_SECS < 256 {
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds % HOUR_SECS == 0 && seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
if seconds.is_multiple_of(HOUR_SECS) && seconds / HOUR_SECS < 256 {
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
// Minutes: truncating division
if seconds / MINUTE_SECS < 256 {
return TTL { count: (seconds / MINUTE_SECS) as u8, unit: TTL_UNIT_MINUTE };
return TTL {
count: (seconds / MINUTE_SECS) as u8,
unit: TTL_UNIT_MINUTE,
};
}
// Second pass: truncating division from smallest to largest
if seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
if seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
TTL::EMPTY
}
+122 -152
View File
@@ -97,12 +97,13 @@ impl NeedleMapMetric {
self.file_byte_count
.fetch_add(new_size.0 as u64, Ordering::Relaxed);
// Go: if oldSize > 0 && oldSize.IsValid() { LogDeletionCounter(oldSize) }
if let Some(old_val) = old {
if old_val.size.0 > 0 && old_val.size.is_valid() {
self.deletion_count.fetch_add(1, Ordering::Relaxed);
self.deletion_byte_count
.fetch_add(old_val.size.0 as u64, Ordering::Relaxed);
}
if let Some(old_val) = old
&& old_val.size.0 > 0
&& old_val.size.is_valid()
{
self.deletion_count.fetch_add(1, Ordering::Relaxed);
self.deletion_byte_count
.fetch_add(old_val.size.0 as u64, Ordering::Relaxed);
}
}
@@ -225,6 +226,12 @@ pub struct CompactNeedleMap {
idx_file_offset: u64,
}
impl Default for CompactNeedleMap {
fn default() -> Self {
Self::new()
}
}
impl CompactNeedleMap {
/// Create a new empty in-memory map.
pub fn new() -> Self {
@@ -465,9 +472,9 @@ impl RedbNeedleMap {
/// loses at most the writes since the last checkpoint from redb, and
/// the next load replays them from .idx.
fn begin_write_no_fsync(db: &Database) -> io::Result<redb::WriteTransaction> {
let mut txn = db.begin_write().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb begin_write: {}", e))
})?;
let mut txn = db
.begin_write()
.map_err(|e| io::Error::other(format!("redb begin_write: {}", e)))?;
let _ = txn.set_durability(Durability::None);
Ok(txn)
}
@@ -501,7 +508,7 @@ impl RedbNeedleMap {
pub fn checkpoint(&mut self, sync_idx: bool) -> io::Result<()> {
let txn = self.begin_checkpoint(sync_idx)?;
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
self.writes_since_checkpoint = 0;
Ok(())
}
@@ -516,17 +523,17 @@ impl RedbNeedleMap {
if sync_idx {
self.sync()?;
}
let mut txn = self.db_or_err()?.begin_write().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb begin_write: {}", e))
})?;
let mut txn = self
.db_or_err()?
.begin_write()
.map_err(|e| io::Error::other(format!("redb begin_write: {}", e)))?;
txn.set_quick_repair(true);
if self.idx_file.is_some() {
let mut meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e))
})?;
meta.insert(META_IDX_SIZE, self.idx_file_offset).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert meta: {}", e))
})?;
let mut meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
meta.insert(META_IDX_SIZE, self.idx_file_offset)
.map_err(|e| io::Error::other(format!("redb insert meta: {}", e)))?;
}
Ok(txn)
}
@@ -538,22 +545,20 @@ impl RedbNeedleMap {
let db = Database::builder()
.set_cache_size(cache_bytes)
.create(db_path)
.map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb create error: {}", e))
})?;
.map_err(|e| io::Error::other(format!("redb create error: {}", e)))?;
// Ensure tables exist
let txn = Self::begin_write_no_fsync(&db)?;
{
let _table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let _meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table meta: {}", e))
})?;
let _table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let _meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table meta: {}", e)))?;
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
Ok(RedbNeedleMap {
db: Some(db),
@@ -572,16 +577,14 @@ impl RedbNeedleMap {
fn save_idx_size_meta(&self, idx_size: u64) -> io::Result<()> {
let txn = Self::begin_write_no_fsync(self.db_or_err()?)?;
{
let mut meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e))
})?;
meta.insert(META_IDX_SIZE, idx_size).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert meta: {}", e))
})?;
let mut meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
meta.insert(META_IDX_SIZE, idx_size)
.map_err(|e| io::Error::other(format!("redb insert meta: {}", e)))?;
}
txn.commit().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb commit meta: {}", e))
})?;
txn.commit()
.map_err(|e| io::Error::other(format!("redb commit meta: {}", e)))?;
Ok(())
}
@@ -590,22 +593,18 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
// experimental-api-5 drops inherent ReadOnlyTable::get ('static guard).
// ReadableTable::get guard borrows `meta`; bind the match so the
// temporary Result is dropped before `meta`.
let result = match meta.get(META_IDX_SIZE) {
// ReadableTable::get guard borrows `meta`; edition 2024 drops the tail
// expression's temporaries before `meta`, so no extra binding is needed.
match meta.get(META_IDX_SIZE) {
Ok(Some(guard)) => Ok(Some(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get meta: {}", e),
)),
};
result
Err(e) => Err(io::Error::other(format!("redb get meta: {}", e))),
}
}
/// Load from an .idx file, reusing an existing .rdb if it is consistent.
@@ -648,7 +647,7 @@ impl RedbNeedleMap {
let db = Database::builder()
.set_cache_size(cache_bytes)
.open(db_path)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open: {}", e)))?;
let mut nm = RedbNeedleMap {
db: Some(db),
@@ -663,14 +662,11 @@ impl RedbNeedleMap {
let stored_idx_size = nm
.read_idx_size_meta()?
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "no idx_size in redb meta"))?;
.ok_or_else(|| io::Error::other("no idx_size in redb meta"))?;
if stored_idx_size > idx_size {
// .idx shrank — corrupted or truncated, need full rebuild
return Err(io::Error::new(
io::ErrorKind::Other,
"idx file smaller than stored size",
));
return Err(io::Error::other("idx file smaller than stored size"));
}
// Counters come from the whole .idx history, never from the table,
@@ -683,40 +679,37 @@ impl RedbNeedleMap {
let start_entry = stored_idx_size / NEEDLE_MAP_ENTRY_SIZE as u64;
let txn = Self::begin_write_no_fsync(nm.db.as_ref().unwrap())?;
{
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
idx::walk_index_file(reader, start_entry, |key, offset, size| {
let key_u64: u64 = key.into();
if offset.is_zero() || size.is_deleted() {
// Delete: store a tombstone (negative size, original
// offset) over a live value; already deleted is a no-op.
if let Ok(Some(old)) = nm.get_via_table(&table, key_u64) {
if old.size.is_valid() {
let deleted_nv = NeedleValue {
offset: old.offset,
size: Size(-(old.size.0)),
};
let packed = pack_needle_value(&deleted_nv);
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
)
})?;
}
if let Ok(Some(old)) = nm.get_via_table(&table, key_u64)
&& old.size.is_valid()
{
let deleted_nv = NeedleValue {
offset: old.offset,
size: Size(-(old.size.0)),
};
let packed = pack_needle_value(&deleted_nv);
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
} else {
let packed = pack_needle_value(&NeedleValue { offset, size });
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert: {}", e))
})?;
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
Ok(())
})?;
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
nm.save_idx_size_meta(idx_size)?;
}
@@ -734,10 +727,7 @@ impl RedbNeedleMap {
match table.get(key_u64) {
Ok(Some(guard)) => Ok(packed_to_needle_value(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
)),
Err(e) => Err(io::Error::other(format!("redb get: {}", e))),
}
}
@@ -790,13 +780,13 @@ impl RedbNeedleMap {
let txn = Self::begin_write_no_fsync(nm.db.as_ref().unwrap())?;
{
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
if !unlinked {
table.retain(|_, _| false).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb retain: {}", e))
})?;
table
.retain(|_, _| false)
.map_err(|e| io::Error::other(format!("redb retain: {}", e)))?;
}
#[cfg(not(feature = "redb-experimental-cursor"))]
@@ -804,9 +794,9 @@ impl RedbNeedleMap {
for (key, nv) in &entries {
let key_u64: u64 = (*key).into();
let packed = pack_needle_value(nv);
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert: {}", e))
})?;
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
}
#[cfg(feature = "redb-experimental-cursor")]
@@ -835,7 +825,7 @@ impl RedbNeedleMap {
}
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
nm.save_idx_size_meta(idx_size)?;
Ok(())
@@ -901,23 +891,16 @@ impl RedbNeedleMap {
Ok(t) => t,
Err(e) => {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb open_table: {}", e),
));
return Err(io::Error::other(format!("redb open_table: {}", e)));
}
};
let result = match table.insert(key_u64, packed.as_slice()) {
match table.insert(key_u64, packed.as_slice()) {
Ok(prev) => prev.and_then(|g| packed_to_needle_value(g.value())),
Err(e) => {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
));
return Err(io::Error::other(format!("redb insert: {}", e)));
}
};
result
}
};
match txn.commit() {
Ok(()) => old,
@@ -925,8 +908,7 @@ impl RedbNeedleMap {
// Transaction rolled back, database still usable:
// truncate the orphan .idx row.
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
return Err(io::Error::other(
"redb commit: Transaction was poisoned by a panic",
));
}
@@ -935,12 +917,9 @@ impl RedbNeedleMap {
// visible and redb refuses further writes. Keep
// the .idx row (do NOT truncate) and reopen from
// .idx to repair redb's internal state.
let err = io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e));
let err = io::Error::other(format!("redb commit: {}", e));
if let Err(reopen_err) = self.reopen_from_idx() {
tracing::warn!(
"redb reopen after put commit error failed: {}",
reopen_err
);
tracing::warn!("redb reopen after put commit error failed: {}", reopen_err);
}
return Err(err);
}
@@ -968,39 +947,32 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
// experimental-api-5 drops inherent ReadOnlyTable::get ('static guard).
// ReadableTable::get guard borrows `table`; bind the match so the
// temporary Result is dropped before `table`.
let result = match table.get(key_u64) {
// ReadableTable::get guard borrows `table`; edition 2024 drops the tail
// expression's temporaries before `table`, so no extra binding is needed.
match table.get(key_u64) {
Ok(Some(guard)) => Ok(packed_to_needle_value(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
)),
};
result
Err(e) => Err(io::Error::other(format!("redb get: {}", e))),
}
}
/// Mark a needle as deleted. Appends tombstone to .idx file, negates size in redb.
pub fn delete(&mut self, key: NeedleId, offset: Offset) -> io::Result<Option<Size>> {
let key_u64: u64 = key.into();
let txn = Self::begin_write_no_fsync(self.db_or_err()?)?;
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let old = match table.get(key_u64) {
Ok(Some(guard)) => packed_to_needle_value(guard.value()),
Ok(None) => None,
Err(e) => {
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
));
return Err(io::Error::other(format!("redb get: {}", e)));
}
};
let Some(old) = old.filter(|nv| nv.size.is_valid()) else {
@@ -1021,10 +993,7 @@ impl RedbNeedleMap {
drop(table);
if let Err(e) = insert_res {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
));
return Err(io::Error::other(format!("redb insert: {}", e)));
}
match txn.commit() {
Ok(()) => {}
@@ -1032,8 +1001,7 @@ impl RedbNeedleMap {
// Transaction rolled back, database still usable:
// truncate the orphan .idx row.
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
return Err(io::Error::other(
"redb commit: Transaction was poisoned by a panic",
));
}
@@ -1042,7 +1010,7 @@ impl RedbNeedleMap {
// and redb refuses further writes. Keep the .idx row
// (do NOT truncate) and reopen from .idx to repair
// redb's internal state.
let err = io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e));
let err = io::Error::other(format!("redb commit: {}", e));
if let Err(reopen_err) = self.reopen_from_idx() {
tracing::warn!(
"redb reopen after delete commit error failed: {}",
@@ -1105,10 +1073,10 @@ impl RedbNeedleMap {
/// after the orphan, `idx_file_offset` advances past it, and a later
/// checkpoint records an offset that makes the reload skip the orphan.
fn truncate_idx_to_offset(&mut self) {
if let Some(ref mut idx_file) = self.idx_file {
if let Err(e) = idx_file.truncate_to(self.idx_file_offset) {
tracing::warn!("failed to truncate orphan .idx row: {}", e);
}
if let Some(ref mut idx_file) = self.idx_file
&& let Err(e) = idx_file.truncate_to(self.idx_file_offset)
{
tracing::warn!("failed to truncate orphan .idx row: {}", e);
}
}
@@ -1198,10 +1166,10 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let mut file = std::fs::OpenOptions::new()
.write(true)
@@ -1212,18 +1180,17 @@ impl RedbNeedleMap {
// redb iterates in key order (u64 ascending)
let iter = table
.iter()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb iter: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb iter: {}", e)))?;
for entry in iter {
let (key_guard, val_guard) = entry.map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb iter next: {}", e))
})?;
let (key_guard, val_guard) =
entry.map_err(|e| io::Error::other(format!("redb iter next: {}", e)))?;
let key_u64: u64 = key_guard.value();
let bytes: &[u8] = val_guard.value();
if let Some(nv) = packed_to_needle_value(bytes) {
if nv.size.is_valid() {
idx::write_index_entry(&mut file, NeedleId(key_u64), nv.offset, nv.size)?;
}
if let Some(nv) = packed_to_needle_value(bytes)
&& nv.size.is_valid()
{
idx::write_index_entry(&mut file, NeedleId(key_u64), nv.offset, nv.size)?;
}
}
file.sync_all()?;
@@ -1682,6 +1649,7 @@ mod tests {
.read(true)
.write(true)
.create(true)
.truncate(false)
.open(&idx_path)
.unwrap();
let idx_size = idx_file.metadata().unwrap().len();
@@ -2344,12 +2312,14 @@ mod tests {
reloaded.deleted_count(),
reloaded.deleted_size(),
);
assert_eq!(
after, live,
"close_first={close_first} rebuild={rebuild}"
);
assert_eq!(after, live, "close_first={close_first} rebuild={rebuild}");
assert_eq!(reloaded.get(NeedleId(1)).unwrap().unwrap().size, Size(200));
assert!(reloaded.get(NeedleId(2)).unwrap().map_or(true, |v| v.size.is_deleted()));
assert!(
reloaded
.get(NeedleId(2))
.unwrap()
.is_none_or(|v| v.size.is_deleted())
);
}
}
}
@@ -31,7 +31,7 @@ struct CompactEntry {
}
impl CompactEntry {
fn to_needle_value(&self) -> NeedleValue {
fn to_needle_value(self) -> NeedleValue {
NeedleValue {
offset: Offset::from_bytes(&self.offset),
size: self.size,
@@ -226,10 +226,7 @@ impl SortedFileNeedleMap {
.fail_sdx_mark
.load(std::sync::atomic::Ordering::Relaxed)
{
return Err(io::Error::new(
io::ErrorKind::Other,
"injected .sdx mark failure",
));
return Err(io::Error::other("injected .sdx mark failure"));
}
let mut buf = [0u8; SIZE_SIZE];
TOMBSTONE_FILE_SIZE.to_bytes(&mut buf);
@@ -309,7 +306,7 @@ impl SortedFileNeedleMap {
let rows = rows_per_read.min(entry_count - done) as usize;
let bytes = &mut block[..rows * NEEDLE_MAP_ENTRY_SIZE];
read_exact_at(&file, bytes, done * NEEDLE_MAP_ENTRY_SIZE as u64)?;
for entry in bytes.chunks_exact(NEEDLE_MAP_ENTRY_SIZE) {
for entry in bytes.as_chunks::<NEEDLE_MAP_ENTRY_SIZE>().0 {
let (key, offset, size) = idx_entry_from_bytes(entry);
if !size.is_valid() || pending.contains_key(&key) {
continue; // deleted in place, or still awaiting that mark
+36 -41
View File
@@ -369,6 +369,7 @@ impl Store {
}
/// Create a new volume, placing it on the location with the most free space.
#[expect(clippy::too_many_arguments)]
pub fn add_volume(
&mut self,
vid: VolumeId,
@@ -383,10 +384,10 @@ impl Store {
return Err(VolumeError::AlreadyExists);
}
let loc_idx = self.find_free_location(&disk_type).ok_or_else(|| {
VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("no free location for disk type {:?}", disk_type),
))
VolumeError::Io(io::Error::other(format!(
"no free location for disk type {:?}",
disk_type
)))
})?;
self.locations[loc_idx].create_volume(
@@ -459,7 +460,7 @@ impl Store {
}
// Find the location where the .dat file exists
for loc in &mut self.locations {
if &loc.disk_type != &disk_type {
if loc.disk_type != disk_type {
continue;
}
let base = crate::storage::volume::volume_file_name(&loc.directory, collection, vid);
@@ -472,10 +473,10 @@ impl Store {
// Fail the mount so the caller (VolumeCopy) treats it as an error.
let note_path = format!("{}.note", base);
if std::path::Path::new(&note_path).exists() {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("volume {} copy incomplete: .note still present", vid),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"volume {} copy incomplete: .note still present",
vid
))));
}
return loc.create_volume(
vid,
@@ -600,7 +601,7 @@ impl Store {
// register a phantom normal volume that shadows the real EC volume.
// Match the guard in load_existing_volumes: only mount when a real
// .dat is present, or the .vif points at a remote-tiered file.
let dat_exists = std::fs::metadata(&format!("{}.dat", base_path))
let dat_exists = std::fs::metadata(format!("{}.dat", base_path))
.map(|m| !m.is_dir())
.unwrap_or(false);
let idx_base = crate::storage::volume::volume_file_name(
@@ -664,13 +665,12 @@ impl Store {
for entry in entries.flatten() {
let name = entry.file_name();
let name = name.to_string_lossy();
if let Some((collection, file_vid)) = parse_volume_filename(&name) {
if file_vid == vid {
if let Some(base) = strip_volume_suffix(&name) {
let base_path = format!("{}/{}", loc.directory, base);
results.push((loc_idx, base_path, collection));
}
}
if let Some((collection, file_vid)) = parse_volume_filename(&name)
&& file_vid == vid
&& let Some(base) = strip_volume_suffix(&name)
{
let base_path = format!("{}/{}", loc.directory, base);
results.push((loc_idx, base_path, collection));
}
}
}
@@ -842,10 +842,8 @@ impl Store {
let vol_count = loc.volumes_len() as i32;
let loc_ec_shards = loc.ec_shard_count();
let ec_equivalent = ((loc_ec_shards
+ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT
- 1)
/ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
let ec_equivalent = loc_ec_shards
.div_ceil(crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
as i32;
let mut max_count = vol_count + ec_equivalent;
@@ -1094,10 +1092,10 @@ impl Store {
/// first disk and miss shards that live on a sibling.
pub fn find_ec_shard_location(&self, vid: VolumeId, shard_id: u32) -> Option<usize> {
for (i, loc) in self.locations.iter().enumerate() {
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
return Some(i);
}
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id as u8)
{
return Some(i);
}
}
None
@@ -1106,16 +1104,12 @@ impl Store {
/// Like [`Self::find_ec_shard_location`] but returns the EcVolume
/// reference directly. Borrows the store immutably for the
/// EcVolume's lifetime.
pub fn find_ec_volume_with_shard(
&self,
vid: VolumeId,
shard_id: u32,
) -> Option<&EcVolume> {
pub fn find_ec_volume_with_shard(&self, vid: VolumeId, shard_id: u32) -> Option<&EcVolume> {
for loc in &self.locations {
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
return Some(ecv);
}
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id as u8)
{
return Some(ecv);
}
}
None
@@ -1144,9 +1138,9 @@ impl Store {
if found_vol.is_none() {
found_vol = Some(ecv);
}
for shard_id in 0..max_shard_count {
if dirs[shard_id].is_none() && ecv.has_shard(shard_id as u8) {
dirs[shard_id] = Some(loc.directory.clone());
for (shard_id, dir) in dirs.iter_mut().enumerate() {
if dir.is_none() && ecv.has_shard(shard_id as u8) {
*dir = Some(loc.directory.clone());
}
}
}
@@ -1516,9 +1510,10 @@ fn load_vif_volume_info(path: &str) -> Result<VifVolumeInfo, VolumeError> {
read_only: bool,
}
if let Ok(legacy) = serde_json::from_str::<LegacyVolumeInfo>(&content) {
let mut vif = VifVolumeInfo::default();
vif.read_only = legacy.read_only;
return Ok(vif);
return Ok(VifVolumeInfo {
read_only: legacy.read_only,
..VifVolumeInfo::default()
});
}
Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
@@ -1528,7 +1523,7 @@ fn load_vif_volume_info(path: &str) -> Result<VifVolumeInfo, VolumeError> {
fn save_vif_volume_info(path: &str, info: &VifVolumeInfo) -> Result<(), VolumeError> {
let content = serde_json::to_string_pretty(info)
.map_err(|e| VolumeError::Io(io::Error::new(io::ErrorKind::Other, e.to_string())))?;
.map_err(|e| VolumeError::Io(io::Error::other(e.to_string())))?;
std::fs::write(path, content)?;
Ok(())
}
@@ -80,6 +80,11 @@ struct EcxOwnerInfo {
idx_dir: String,
}
/// One unit of reconcile work: the disk holding orphan shards, the volume
/// they belong to, the shard files, the `.ecx` owner, and whether the
/// mirror already installed sidecars locally (`use_local_idx`).
type OrphanShardLoad = (usize, EcKey, Vec<(String, u32)>, EcxOwnerInfo, bool);
impl Store {
/// Run cross-disk orphan-shard reconciliation. Should be called
/// after every DiskLocation has finished its per-disk EC scan.
@@ -98,7 +103,7 @@ impl Store {
// `use_local_idx` is the post-mirror fast path: when the
// mirror already installed sidecars locally, mount against
// loc.idx_directory instead of the owner disk.
let mut to_load: Vec<(usize, EcKey, Vec<(String, u32)>, EcxOwnerInfo, bool)> = Vec::new();
let mut to_load: Vec<OrphanShardLoad> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
let orphans = collect_orphan_ec_shards(loc, loc_idx);
for (key, shards) in orphans {
@@ -293,10 +298,10 @@ impl Store {
// may be sole copies of a distributed volume.
let mut node_wide_bits = ev.shard_bits().0;
for other in &self.locations {
if let Some(other_ev) = other.find_ec_volume(*vid) {
if other_ev.collection == ev.collection {
node_wide_bits |= other_ev.shard_bits().0;
}
if let Some(other_ev) = other.find_ec_volume(*vid)
&& other_ev.collection == ev.collection
{
node_wide_bits |= other_ev.shard_bits().0;
}
}
let node_wide = node_wide_bits.count_ones() as usize;
@@ -499,6 +504,53 @@ impl Store {
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, u32)>> {
let mut orphans: HashMap<EcKey, Vec<(String, u32)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id as u8)
{
continue;
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
#[cfg(test)]
mod tests {
use super::*;
@@ -1759,50 +1811,3 @@ mod tests {
assert!(std::path::Path::new(&format!("{}.ecx", ec_base)).exists());
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, u32)>> {
let mut orphans: HashMap<EcKey, Vec<(String, u32)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
continue;
}
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
+4 -24
View File
@@ -155,7 +155,7 @@ impl Size {
return 0;
}
if self.0 < 0 {
return (self.0 * -1) as u32;
return -self.0 as u32;
}
self.0 as u32
}
@@ -284,8 +284,9 @@ impl fmt::Display for Offset {
// DiskType
// ============================================================================
#[derive(Debug, Clone, PartialEq, Eq, Hash)]
#[derive(Debug, Clone, PartialEq, Eq, Hash, Default)]
pub enum DiskType {
#[default]
HardDrive,
Ssd,
Custom(String),
@@ -319,12 +320,6 @@ impl fmt::Display for DiskType {
}
}
impl Default for DiskType {
fn default() -> Self {
DiskType::HardDrive
}
}
// ============================================================================
// VolumeId
// ============================================================================
@@ -397,7 +392,7 @@ impl From<u8> for Version {
///
/// Fields are split into request-side options (set by the caller) and response-side
/// flags (set during the read to communicate status back).
#[derive(Debug, Clone)]
#[derive(Debug, Clone, Default)]
pub struct ReadOption {
// -- request --
/// If true, allow reading needles that have been soft-deleted.
@@ -423,21 +418,6 @@ pub struct ReadOption {
pub read_buffer_size: i32,
}
impl Default for ReadOption {
fn default() -> Self {
ReadOption {
read_deleted: false,
attempt_meta_only: false,
must_meta_only: false,
is_meta_only: false,
volume_revision: 0,
is_out_of_range: false,
has_slow_read: false,
read_buffer_size: 0,
}
}
}
// ============================================================================
// NeedleMapEntry helpers (for .idx file)
// ============================================================================
+245 -277
View File
@@ -106,7 +106,7 @@ fn exceeds_expected_compacted_size(expected_live_bytes: u64, dst_dat_size: u64)
pub fn is_storage_io_error(e: &io::Error) -> bool {
#[cfg(unix)]
{
return e.raw_os_error() == Some(libc::EIO);
e.raw_os_error() == Some(libc::EIO)
}
#[cfg(windows)]
{
@@ -249,7 +249,7 @@ struct OldVersionVifVolumeInfo {
impl OldVersionVifVolumeInfo {
/// Convert to the standard VifVolumeInfo, mapping destroy_time -> expire_at_sec.
fn to_vif(self) -> VifVolumeInfo {
fn into_vif(self) -> VifVolumeInfo {
VifVolumeInfo {
files: self.files,
version: self.version,
@@ -511,7 +511,7 @@ impl RemoteDatFile {
let data = self
.backend
.read_range_blocking(&self.key, offset, buf.len())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e))?;
.map_err(io::Error::other)?;
if data.len() != buf.len() {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -532,6 +532,10 @@ impl RemoteDatFile {
// Volume
// ============================================================================
/// One raw needle as `scan_raw_needles_from` yields it: the header bytes,
/// the body bytes, and the needle's `append_at_ns`.
pub type RawNeedleEntry = (Vec<u8>, Vec<u8>, u64);
pub struct Volume {
pub id: VolumeId,
dir: String,
@@ -608,6 +612,7 @@ fn read_exact_at(file: &File, buf: &mut [u8], mut offset: u64) -> io::Result<()>
impl Volume {
/// Create and load a volume from disk.
#[expect(clippy::too_many_arguments)]
pub fn new(
dirname: &str,
dir_idx: &str,
@@ -899,18 +904,18 @@ impl Volume {
// so vacuum doesn't silently drop reachable data based on a
// corrupt .idx left over from a crashed batched write.
// See issue #8928.
if let Some(ref nm) = self.nm {
if let Ok(dat_size) = self.current_dat_file_size() {
let max_end = nm.max_needle_end();
if dat_size > 0 && max_end > dat_size as i64 {
self.no_write_or_delete = true;
warn!(
volume_id = self.id.0,
max_needle_end = max_end,
dat_size,
"idx references bytes past end of .dat; marking volume read-only"
);
}
if let Some(ref nm) = self.nm
&& let Ok(dat_size) = self.current_dat_file_size()
{
let max_end = nm.max_needle_end();
if dat_size > 0 && max_end > dat_size as i64 {
self.no_write_or_delete = true;
warn!(
volume_id = self.id.0,
max_needle_end = max_end,
dat_size,
"idx references bytes past end of .dat; marking volume read-only"
);
}
}
}
@@ -1071,10 +1076,7 @@ impl Volume {
let mut nm = CompactNeedleMap::load_from_idx(&mut idx_reader, self.version())?;
// Re-open for append-only writes
let write_file = OpenOptions::new()
.write(true)
.append(true)
.open(&idx_path)?;
let write_file = OpenOptions::new().append(true).open(idx_path)?;
nm.set_idx_file(Box::new(write_file), idx_size);
self.nm = Some(NeedleMap::InMemory(nm));
}
@@ -1132,10 +1134,7 @@ impl Volume {
)?;
// Re-open for append-only writes
let write_file = OpenOptions::new()
.write(true)
.append(true)
.open(&idx_path)?;
let write_file = OpenOptions::new().append(true).open(idx_path)?;
nm.set_idx_file(Box::new(write_file), idx_size);
self.nm = Some(NeedleMap::Redb(nm));
}
@@ -1208,10 +1207,7 @@ impl Volume {
remote_dat_file.read_exact_at(buf, offset)?;
Ok(())
} else {
Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
"dat file not open",
)))
Err(VolumeError::Io(io::Error::other("dat file not open")))
}
}
@@ -1283,9 +1279,10 @@ impl Volume {
}
fn maybe_write_super_block(&mut self, version: Version) -> Result<(), VolumeError> {
let dat_file = self.dat_file.as_mut().ok_or_else(|| {
VolumeError::Io(io::Error::new(io::ErrorKind::Other, "dat file not open"))
})?;
let dat_file = self
.dat_file
.as_mut()
.ok_or_else(|| VolumeError::Io(io::Error::other("dat file not open")))?;
let dat_size = dat_file.metadata()?.len();
if dat_size == 0 {
@@ -1353,18 +1350,18 @@ impl Volume {
}
// TTL expiry check
if n.has_ttl() {
if let Some(ref ttl) = n.ttl {
let ttl_minutes = ttl.minutes();
if ttl_minutes > 0 && n.has_last_modified_date() {
let expire_at_ns = n.append_at_ns + (ttl_minutes as u64) * 60 * 1_000_000_000;
let now_ns = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_nanos() as u64;
if now_ns >= expire_at_ns {
return Err(VolumeError::NotFound);
}
if n.has_ttl()
&& let Some(ref ttl) = n.ttl
{
let ttl_minutes = ttl.minutes();
if ttl_minutes > 0 && n.has_last_modified_date() {
let expire_at_ns = n.append_at_ns + (ttl_minutes as u64) * 60 * 1_000_000_000;
let now_ns = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_nanos() as u64;
if now_ns >= expire_at_ns {
return Err(VolumeError::NotFound);
}
}
}
@@ -1417,7 +1414,7 @@ impl Volume {
let mut buf = vec![0u8; actual_size as usize];
self.read_exact_at_backend(&mut buf, offset as u64)?;
n.read_bytes(&mut buf, offset, size, version)?;
n.read_bytes(&buf, offset, size, version)?;
Ok(())
}
@@ -1505,7 +1502,7 @@ impl Volume {
if size.0 == 0 || version == VERSION_1 {
// Tombstone or V1: no body data section, tail starts right after header
let meta_size = actual_size - NEEDLE_HEADER_SIZE as i64;
if meta_size < 0 || meta_size > 128 * 1024 {
if !(0..=128 * 1024).contains(&meta_size) {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
@@ -1536,7 +1533,7 @@ impl Volume {
let meta_size = stop_offset - start_offset;
// Sanity check: reject metadata sizes > 128KB (matching Go's ReadNeedleMeta guard)
if meta_size < 0 || meta_size > 128 * 1024 {
if !(0..=128 * 1024).contains(&meta_size) {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
@@ -1595,7 +1592,7 @@ impl Volume {
let mut read_and_parse = |off: i64| -> Result<(), VolumeError> {
let mut buf = vec![0u8; actual_size as usize];
self.read_exact_at_backend(&mut buf, off as u64)?;
n.read_bytes_meta_only(&mut buf, off, read_size, version)?;
n.read_bytes_meta_only(&buf, off, read_size, version)?;
Ok(())
};
@@ -1612,18 +1609,18 @@ impl Volume {
}
// TTL expiry check
if n.has_ttl() {
if let Some(ref ttl) = n.ttl {
let ttl_minutes = ttl.minutes();
if ttl_minutes > 0 && n.has_last_modified_date() {
let expire_at_ns = n.append_at_ns + (ttl_minutes as u64) * 60 * 1_000_000_000;
let now_ns = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_nanos() as u64;
if now_ns >= expire_at_ns {
return Err(VolumeError::NotFound);
}
if n.has_ttl()
&& let Some(ref ttl) = n.ttl
{
let ttl_minutes = ttl.minutes();
if ttl_minutes > 0 && n.has_last_modified_date() {
let expire_at_ns = n.append_at_ns + (ttl_minutes as u64) * 60 * 1_000_000_000;
let now_ns = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_nanos() as u64;
if now_ns >= expire_at_ns {
return Err(VolumeError::NotFound);
}
}
}
@@ -1718,14 +1715,11 @@ impl Volume {
fn flush_dat(&self) -> io::Result<()> {
#[cfg(test)]
if self.fail_fsync_for_test {
return Err(io::Error::new(
io::ErrorKind::Other,
"injected fsync failure",
));
return Err(io::Error::other("injected fsync failure"));
}
match self.dat_file.as_ref() {
Some(dat_file) => dat_file.sync_all(),
None => Err(io::Error::new(io::ErrorKind::Other, "dat file not open")),
None => Err(io::Error::other("dat file not open")),
}
}
@@ -1743,7 +1737,7 @@ impl Volume {
fn flush_idx(&mut self) -> Result<(), VolumeError> {
#[cfg(test)]
if self.fail_idx_sync_for_test {
let e = io::Error::new(io::ErrorKind::Other, "injected idx sync failure");
let e = io::Error::other("injected idx sync failure");
self.no_write_or_delete = true;
return Err(VolumeError::Io(e));
}
@@ -1804,18 +1798,18 @@ impl Volume {
}
// Cookie validation for existing needle (matches Go: check whenever nm.Get returns ok)
if let Some(nm) = &self.nm {
if let Some(nv) = nm.get(n.id)? {
let mut existing = Needle::default();
// Read only the header to check cookie
self.read_needle_header_unlocked(&mut existing, nv.offset.to_actual_offset())?;
if let Some(nm) = &self.nm
&& let Some(nv) = nm.get(n.id)?
{
let mut existing = Needle::default();
// Read only the header to check cookie
self.read_needle_header_unlocked(&mut existing, nv.offset.to_actual_offset())?;
if n.cookie.0 == 0 && !check_cookie {
n.cookie = existing.cookie;
}
if existing.cookie != n.cookie {
return Err(VolumeError::CookieMismatch(n.cookie.0));
}
if n.cookie.0 == 0 && !check_cookie {
n.cookie = existing.cookie;
}
if existing.cookie != n.cookie {
return Err(VolumeError::CookieMismatch(n.cookie.0));
}
}
@@ -1828,28 +1822,26 @@ impl Volume {
// Nothing is published until the bytes are down: an index entry for an
// unflushed append would resolve past the end of the file after a crash,
// and undoing it afterwards would double-count the volume's metrics.
if fsync {
if let Err(e) = self.flush_dat() {
self.check_read_write_error(Some(&e));
let truncated = match self.dat_file.as_ref() {
Some(dat_file) => dat_file.set_len(offset),
None => Ok(()),
};
if let Err(te) = truncated {
// The rejected record is still on the end. A later append
// would bury it mid-file, where the .dat tail check cannot
// see it, so stop taking writes instead.
self.no_write_or_delete = true;
tracing::error!(
"volume {}: failed to truncate back to {} after a failed fsync, \
if fsync && let Err(e) = self.flush_dat() {
self.check_read_write_error(Some(&e));
let truncated = match self.dat_file.as_ref() {
Some(dat_file) => dat_file.set_len(offset),
None => Ok(()),
};
if let Err(te) = truncated {
// The rejected record is still on the end. A later append
// would bury it mid-file, where the .dat tail check cannot
// see it, so stop taking writes instead.
self.no_write_or_delete = true;
tracing::error!(
"volume {}: failed to truncate back to {} after a failed fsync, \
marking read only: {}",
self.id.0,
offset,
te
);
}
return Err(VolumeError::Io(e));
self.id.0,
offset,
te
);
}
return Err(VolumeError::Io(e));
}
self.last_append_at_ns = n.append_at_ns;
@@ -1986,26 +1978,25 @@ impl Volume {
return None;
}
};
if let Some(nv) = existing {
if !nv.offset.is_zero() && nv.size.is_valid() {
let mut old = Needle::default();
let mut ro = ReadOption::default();
if self
.read_needle_data_at_unlocked(
&mut old,
nv.offset.to_actual_offset(),
nv.size,
&mut ro,
)
.is_ok()
{
if old.cookie == n.cookie
&& old.checksum == n.checksum
&& old.data == n.data
{
return Some(old.data_size);
}
}
if let Some(nv) = existing
&& !nv.offset.is_zero()
&& nv.size.is_valid()
{
let mut old = Needle::default();
let mut ro = ReadOption::default();
if self
.read_needle_data_at_unlocked(
&mut old,
nv.offset.to_actual_offset(),
nv.size,
&mut ro,
)
.is_ok()
&& old.cookie == n.cookie
&& old.checksum == n.checksum
&& old.data == n.data
{
return Some(old.data_size);
}
}
}
@@ -2018,9 +2009,10 @@ impl Volume {
let bytes = n.write_bytes(version);
let actual_size = bytes.len() as i64;
let dat_file = self.dat_file.as_mut().ok_or_else(|| {
VolumeError::Io(io::Error::new(io::ErrorKind::Other, "dat file not open"))
})?;
let dat_file = self
.dat_file
.as_mut()
.ok_or_else(|| VolumeError::Io(io::Error::other("dat file not open")))?;
let offset = dat_file.seek(SeekFrom::End(0))?;
@@ -2228,7 +2220,7 @@ impl Volume {
}
let body_length = needle::needle_body_length(size, version);
let total_size = NEEDLE_HEADER_SIZE as i64 + body_length as i64;
let total_size = NEEDLE_HEADER_SIZE as i64 + body_length;
if size.is_deleted() || size.0 <= 0 {
offset += total_size;
@@ -2380,7 +2372,7 @@ impl Volume {
let entries = &mut block[..(end - start) as usize];
idx_file.seek(SeekFrom::Start(start as u64))?;
idx_file.read_exact(entries)?;
for entry in entries.chunks_exact(NEEDLE_MAP_ENTRY_SIZE).rev() {
for entry in entries.as_chunks::<NEEDLE_MAP_ENTRY_SIZE>().0.iter().rev() {
let (key, offset, size) = idx_entry_from_bytes(entry);
if offset.is_zero() || size.is_deleted() {
continue;
@@ -2455,16 +2447,16 @@ impl Volume {
idx_size: i64,
version: Version,
) -> Result<i64, VolumeError> {
if let Ok(dat_size) = self.dat_file_size() {
if dat_size > 0 {
let last_pos = idx_size - NEEDLE_MAP_ENTRY_SIZE as i64;
let mut buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
idx_file.seek(SeekFrom::Start(last_pos as u64))?;
idx_file.read_exact(&mut buf)?;
let (_, offset, size) = idx_entry_from_bytes(&buf);
if !offset.is_zero() && needle_disk_end(offset, size, version) == dat_size as i64 {
return Ok(last_pos);
}
if let Ok(dat_size) = self.dat_file_size()
&& dat_size > 0
{
let last_pos = idx_size - NEEDLE_MAP_ENTRY_SIZE as i64;
let mut buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
idx_file.seek(SeekFrom::Start(last_pos as u64))?;
idx_file.read_exact(&mut buf)?;
let (_, offset, size) = idx_entry_from_bytes(&buf);
if !offset.is_zero() && needle_disk_end(offset, size, version) == dat_size as i64 {
return Ok(last_pos);
}
}
@@ -2753,12 +2745,12 @@ impl Volume {
}
/// Scan raw needle entries from the .dat file starting at `from_offset`.
/// Returns (needle_header_bytes, needle_body_bytes, append_at_ns) for each needle.
/// Returns a [`RawNeedleEntry`] for each needle.
/// Used by VolumeTailSender to stream raw bytes.
pub fn scan_raw_needles_from(
&self,
from_offset: u64,
) -> Result<Vec<(Vec<u8>, Vec<u8>, u64)>, VolumeError> {
) -> Result<Vec<RawNeedleEntry>, VolumeError> {
let version = self.version();
let dat_size = self.current_dat_file_size()?;
let mut entries = Vec::new();
@@ -2770,7 +2762,7 @@ impl Volume {
match self.read_exact_at_backend(&mut header, offset) {
Ok(()) => {}
Err(VolumeError::Io(e)) if e.kind() == io::ErrorKind::UnexpectedEof => break,
Err(e) => return Err(e.into()),
Err(e) => return Err(e),
}
let (_cookie, _id, size) = Needle::parse_header(&header);
@@ -2789,7 +2781,7 @@ impl Volume {
match self.read_exact_at_backend(&mut body, offset + NEEDLE_HEADER_SIZE as u64) {
Ok(()) => {}
Err(VolumeError::Io(e)) if e.kind() == io::ErrorKind::UnexpectedEof => break,
Err(e) => return Err(e.into()),
Err(e) => return Err(e),
}
// Parse the needle to get append_at_ns
@@ -2861,7 +2853,6 @@ impl Volume {
if needs_idx_writer {
let idx_path = self.file_name(".idx");
let write_file = OpenOptions::new()
.write(true)
.append(true)
.create(true)
.open(&idx_path)?;
@@ -3016,7 +3007,7 @@ impl Volume {
// Fall back to OldVersionVolumeInfo (Go's tryOldVersionVolumeInfo):
// maps DestroyTime -> expire_at_sec
if let Ok(old_info) = serde_json::from_str::<OldVersionVifVolumeInfo>(&content) {
let vif_info = old_info.to_vif();
let vif_info = old_info.into_vif();
let pb_info = vif_info.to_pb();
if pb_info.read_only {
self.no_write_or_delete = true;
@@ -3085,7 +3076,7 @@ impl Volume {
}
let content = serde_json::to_string_pretty(&vif)
.map_err(|e| VolumeError::Io(io::Error::new(io::ErrorKind::Other, e.to_string())))?;
.map_err(|e| VolumeError::Io(io::Error::other(e.to_string())))?;
fs::write(&vif_path, content)?;
Ok(())
}
@@ -3106,7 +3097,7 @@ impl Volume {
let vif = VifVolumeInfo::from_pb(&self.volume_info);
let content = serde_json::to_string_pretty(&vif)
.map_err(|e| VolumeError::Io(io::Error::new(io::ErrorKind::Other, e.to_string())))?;
.map_err(|e| VolumeError::Io(io::Error::other(e.to_string())))?;
// fsync the .vif so a tiered volume's remote reference is durable before the
// caller acts on it, e.g. deletes the remote object (matches Go util.WriteFile).
let mut f = OpenOptions::new()
@@ -3160,9 +3151,10 @@ impl Volume {
pub fn set_replica_placement(&mut self, rp: ReplicaPlacement) -> Result<(), VolumeError> {
self.super_block.replica_placement = rp;
let bytes = self.super_block.to_bytes();
let dat_file = self.dat_file.as_mut().ok_or_else(|| {
VolumeError::Io(io::Error::new(io::ErrorKind::Other, "dat file not open"))
})?;
let dat_file = self
.dat_file
.as_mut()
.ok_or_else(|| VolumeError::Io(io::Error::other("dat file not open")))?;
dat_file.seek(SeekFrom::Start(0))?;
dat_file.write_all(&bytes)?;
dat_file.sync_all()?;
@@ -3213,7 +3205,7 @@ impl Volume {
self.read_exact_at_backend(&mut buf, actual_offset)?;
let mut n = Needle::default();
n.read_bytes_meta_only(&mut buf, offset.to_actual_offset(), size, version)?;
n.read_bytes_meta_only(&buf, offset.to_actual_offset(), size, version)?;
Ok(n.append_at_ns)
}
@@ -3322,9 +3314,10 @@ impl Volume {
if self.is_read_only() {
return Err(VolumeError::ReadOnly);
}
let dat_file = self.dat_file.as_mut().ok_or_else(|| {
VolumeError::Io(io::Error::new(io::ErrorKind::Other, "dat file not open"))
})?;
let dat_file = self
.dat_file
.as_mut()
.ok_or_else(|| VolumeError::Io(io::Error::other("dat file not open")))?;
dat_file.seek(SeekFrom::Start(offset as u64))?;
dat_file.write_all(needle_blob)?;
Ok(())
@@ -3368,36 +3361,33 @@ impl Volume {
// Dedup check: if the same needle already exists with matching content, skip the write.
// Matches Go's WriteNeedleBlob which reads existing needle and compares cookie+checksum+data.
if let Some(nm) = &self.nm {
if let Some(nv) = nm.get(needle_id)? {
if nv.size == size {
let version = self.version();
// Read existing needle from disk
let mut old_needle = Needle::default();
let mut ro = ReadOption::default();
if self
.read_needle_data_at_unlocked(
&mut old_needle,
nv.offset.to_actual_offset(),
nv.size,
&mut ro,
)
.is_ok()
{
// Parse the incoming blob into a needle
let mut new_needle = Needle::default();
if new_needle
.read_bytes(needle_blob, nv.offset.to_actual_offset(), size, version)
.is_ok()
{
if old_needle.cookie == new_needle.cookie
&& old_needle.checksum == new_needle.checksum
&& old_needle.data == new_needle.data
{
return Ok(());
}
}
}
if let Some(nm) = &self.nm
&& let Some(nv) = nm.get(needle_id)?
&& nv.size == size
{
let version = self.version();
// Read existing needle from disk
let mut old_needle = Needle::default();
let mut ro = ReadOption::default();
if self
.read_needle_data_at_unlocked(
&mut old_needle,
nv.offset.to_actual_offset(),
nv.size,
&mut ro,
)
.is_ok()
{
// Parse the incoming blob into a needle
let mut new_needle = Needle::default();
if new_needle
.read_bytes(needle_blob, nv.offset.to_actual_offset(), size, version)
.is_ok()
&& old_needle.cookie == new_needle.cookie
&& old_needle.checksum == new_needle.checksum
&& old_needle.data == new_needle.data
{
return Ok(());
}
}
}
@@ -3405,13 +3395,10 @@ impl Volume {
// Check volume size limit
let content_size = self.content_size();
if MAX_POSSIBLE_VOLUME_SIZE < content_size + needle_blob.len() as u64 {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!(
"volume size limit {} exceeded! current size is {}",
MAX_POSSIBLE_VOLUME_SIZE, content_size
),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"volume size limit {} exceeded! current size is {}",
MAX_POSSIBLE_VOLUME_SIZE, content_size
))));
}
// Compute monotonic appendAtNs (matches Go: needle.GetAppendAtNs(v.lastAppendAtNs))
@@ -3547,10 +3534,10 @@ impl Volume {
{
// Guard against nil needle map (matches Go's nil check before compaction sync)
if self.nm.is_none() {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("volume {} needle map is nil", self.id),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"volume {} needle map is nil",
self.id
))));
}
// Record state before compaction for makeupDiff
@@ -3630,10 +3617,10 @@ impl Volume {
// compacting away data that might come back on retry.
// See issue #8928.
if !is_skippable_needle_read_error(&e) {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("cannot hydrate needle from file: {}", e),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"cannot hydrate needle from file: {}",
e
))));
}
skipped_needles += 1;
if size.is_valid() {
@@ -3893,13 +3880,10 @@ impl Volume {
// are the only inputs reconcile can roll forward to, so removing them
// mid-commit would strand a decided swap.
if Path::new(&self.file_name(".cpc")).exists() {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!(
"volume {}: refusing cleanup while commit marker present",
self.id
),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"volume {}: refusing cleanup while commit marker present",
self.id
))));
}
let cpd_path = self.file_name(".cpd");
@@ -3914,10 +3898,10 @@ impl Volume {
// Ignore NotFound errors
for e in [e1, e2, e3, e4] {
if let Err(e) = e {
if e.kind() != io::ErrorKind::NotFound {
return Err(e.into());
}
if let Err(e) = e
&& e.kind() != io::ErrorKind::NotFound
{
return Err(e.into());
}
}
@@ -3933,13 +3917,10 @@ impl Volume {
let old_super_block = &self.super_block;
if old_super_block.compaction_revision != self.last_compact_revision {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!(
"current old dat file's compact revision {} is not the expected one {}",
old_super_block.compaction_revision, self.last_compact_revision
),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"current old dat file's compact revision {} is not the expected one {}",
old_super_block.compaction_revision, self.last_compact_revision
))));
}
// Read the new .cpd file's super block and verify its compaction revision is old + 1
@@ -3951,13 +3932,10 @@ impl Volume {
let old_compact_revision = old_super_block.compaction_revision;
let new_compact_revision = new_super_block.compaction_revision;
if old_compact_revision + 1 != new_compact_revision {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!(
"old dat file's compact revision {} + 1 does not equal new dat file's compact revision {}",
old_compact_revision, new_compact_revision
),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"old dat file's compact revision {} + 1 does not equal new dat file's compact revision {}",
old_compact_revision, new_compact_revision
))));
}
let old_idx_path = self.file_name(".idx");
@@ -3984,10 +3962,7 @@ impl Volume {
let cpx_path = self.file_name(".cpx");
let mut dst_dat = OpenOptions::new().read(true).write(true).open(&cpd_path)?;
let mut dst_idx = OpenOptions::new()
.write(true)
.append(true)
.open(&cpx_path)?;
let mut dst_idx = OpenOptions::new().append(true).open(&cpx_path)?;
let mut dat_offset = dst_dat.seek(SeekFrom::End(0))?;
let padding_rem = dat_offset % NEEDLE_PADDING_SIZE as u64;
@@ -4140,10 +4115,10 @@ impl Volume {
self.id, reopen
);
}
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("relocate index for volume {}: move .idx: {e}", self.id),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"relocate index for volume {}: move .idx: {e}",
self.id
))));
}
// The .sdx is a derived sorted index; move it when present, but a
@@ -4173,10 +4148,10 @@ impl Volume {
return Err(VolumeError::NotEmpty);
}
if self.is_compacting {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("volume {} is compacting", self.id),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"volume {} is compacting",
self.id
))));
}
let (storage_name, storage_key) = self.remote_storage_name_key();
@@ -4240,26 +4215,25 @@ impl Volume {
/// EIO error. Matches Go's `checkReadWriteError` in volume_write.go.
fn check_read_write_error(&self, err: Option<&io::Error>) {
use std::sync::atomic::Ordering;
if let Some(e) = err {
if is_storage_io_error(e) {
self.io_error_count.fetch_add(1, Ordering::Relaxed);
if let Ok(mut guard) = self.last_io_error.lock() {
*guard = Some(e.to_string());
}
crate::metrics::STORAGE_IO_ERROR_COUNTER.inc();
return;
if let Some(e) = err
&& is_storage_io_error(e)
{
self.io_error_count.fetch_add(1, Ordering::Relaxed);
if let Ok(mut guard) = self.last_io_error.lock() {
*guard = Some(e.to_string());
}
crate::metrics::STORAGE_IO_ERROR_COUNTER.inc();
return;
}
self.io_error_count.store(0, Ordering::Relaxed);
if let Ok(mut guard) = self.last_io_error.lock() {
if guard.is_some() {
*guard = None;
}
if let Ok(mut guard) = self.last_io_error.lock()
&& guard.is_some()
{
*guard = None;
}
}
/// Returns the last recorded I/O error string, if any.
#[allow(dead_code)]
pub fn last_io_error(&self) -> Option<String> {
self.last_io_error.lock().ok()?.clone()
}
@@ -4427,10 +4401,10 @@ pub(crate) fn fsync_dir(path: &str) -> io::Result<()> {
}
#[cfg(not(windows))]
{
if let Some(parent) = Path::new(path).parent() {
if let Ok(d) = File::open(parent) {
return d.sync_all();
}
if let Some(parent) = Path::new(path).parent()
&& let Ok(d) = File::open(parent)
{
return d.sync_all();
}
Ok(())
}
@@ -4592,34 +4566,33 @@ mod tests {
if let Some(range) = headers
.get(header::RANGE)
.and_then(|value| value.to_str().ok())
&& let Some(spec) = range.strip_prefix("bytes=")
{
if let Some(spec) = range.strip_prefix("bytes=") {
let (start, end) = spec.split_once('-').unwrap();
let start = start.parse::<usize>().unwrap();
let end = if end.is_empty() {
bytes.len().saturating_sub(1)
} else {
end.parse::<usize>().unwrap()
}
.min(bytes.len().saturating_sub(1));
let chunk = bytes[start..=end].to_vec();
let mut response_headers = HeaderMap::new();
response_headers.insert(
header::CONTENT_LENGTH,
HeaderValue::from_str(&chunk.len().to_string()).unwrap(),
);
response_headers.insert(
header::CONTENT_RANGE,
HeaderValue::from_str(&format!(
"bytes {}-{}/{}",
start,
end,
bytes.len()
))
.unwrap(),
);
return (StatusCode::PARTIAL_CONTENT, response_headers, chunk);
let (start, end) = spec.split_once('-').unwrap();
let start = start.parse::<usize>().unwrap();
let end = if end.is_empty() {
bytes.len().saturating_sub(1)
} else {
end.parse::<usize>().unwrap()
}
.min(bytes.len().saturating_sub(1));
let chunk = bytes[start..=end].to_vec();
let mut response_headers = HeaderMap::new();
response_headers.insert(
header::CONTENT_LENGTH,
HeaderValue::from_str(&chunk.len().to_string()).unwrap(),
);
response_headers.insert(
header::CONTENT_RANGE,
HeaderValue::from_str(&format!(
"bytes {}-{}/{}",
start,
end,
bytes.len()
))
.unwrap(),
);
return (StatusCode::PARTIAL_CONTENT, response_headers, chunk);
}
let mut response_headers = HeaderMap::new();
@@ -6485,13 +6458,8 @@ mod tests {
// max_needle_end past dat_size — which is exactly the signal
// volume.load uses to mark the volume read-only.
let bad_offset = Offset::from_actual_offset(dat_size + 4 * 1024 * 1024);
let mut idx_append = OpenOptions::new()
.write(true)
.append(true)
.open(&idx_path)
.unwrap();
idx::write_index_entry(&mut idx_append, NeedleId(9999), bad_offset, Size(1024))
.unwrap();
let mut idx_append = OpenOptions::new().append(true).open(&idx_path).unwrap();
idx::write_index_entry(&mut idx_append, NeedleId(9999), bad_offset, Size(1024)).unwrap();
idx_append.sync_all().unwrap();
let mut idx_reread = File::open(&idx_path).unwrap();
@@ -73,8 +73,10 @@ mod tests {
let empty = master_pb::VolumeInformationMessage::default();
assert_eq!(report_hash(&empty), 10988706248825469653);
let mut one = master_pb::VolumeInformationMessage::default();
one.id = 1;
let one = master_pb::VolumeInformationMessage {
id: 1,
..Default::default()
};
assert_eq!(report_hash(&one), 2035849960016744285);
let full = master_pb::VolumeInformationMessage {
+1 -1
View File
@@ -66,7 +66,7 @@ fn parse_go_version_number() -> Option<String> {
}
}
match (major, minor) {
(Some(maj), Some(min)) => Some(format!("{}.{}", maj, format!("{:02}", min))),
(Some(maj), Some(min)) => Some(format!("{}.{:02}", maj, min)),
_ => None,
}
}
+8
View File
@@ -16,6 +16,14 @@ edition = "2024"
# lance's `aws` feature pulls in).
rust-version = "1.94.1"
[workspace.lints.clippy]
# Every RPC path returns tonic::Status (176 bytes). Boxing it would change
# every handler signature for no gain, so the large-Err lint is off.
result_large_err = "allow"
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
[workspace.dependencies]
anyhow = "1"
async-trait = "0.1"
+3
View File
@@ -26,3 +26,6 @@ tonic-build.workspace = true
# install, and so the version is pinned rather than whatever the platform's
# package manager happens to carry. The same crate seaweed-volume uses.
protoc-bin-vendored = "3"
[lints]
workspace = true
+4 -4
View File
@@ -14,10 +14,10 @@ pub fn server_to_grpc_address(server: &str) -> Option<String> {
let (host, port_part) = server.rsplit_once(':')?;
// "port.grpcPort" states the gRPC port outright.
if let Some((_, grpc_port)) = port_part.split_once('.') {
if let Ok(port) = grpc_port.parse::<u16>() {
return Some(join_host_port(host, port));
}
if let Some((_, grpc_port)) = port_part.split_once('.')
&& let Ok(port) = grpc_port.parse::<u16>()
{
return Some(join_host_port(host, port));
}
let port: u16 = port_part.parse().ok()?;
+3
View File
@@ -16,6 +16,9 @@ pub mod stream;
/// Generated plugin.proto types.
pub mod pb {
// prost gives every oneof its own enum; the variant sizes are the
// messages' own, not a choice made here.
#![allow(clippy::large_enum_variant)]
tonic::include_proto!("plugin");
}
+3
View File
@@ -50,3 +50,6 @@ arrow-array = "58"
arrow-schema = "58"
arrow-cast = "58"
lance-linalg = "10"
[lints]
workspace = true
+3
View File
@@ -11,3 +11,6 @@ name = "seaweed_worker_sort"
[dependencies]
seaweed-worker-core = { path = "../core" }
anyhow.workspace = true
[lints]
workspace = true
+4 -4
View File
@@ -152,10 +152,10 @@ fn parse_field(entry: &str) -> Result<SortField> {
/// back: sorting by the worker's default order instead of the one the table
/// asked for would silently rewrite the table the wrong way.
pub fn resolve(declared: Option<&str>, configured: &str) -> Result<Option<SortSpec>> {
if let Some(declared) = declared {
if let Some(spec) = SortSpec::parse(declared).context("read the table's declared order")? {
return Ok(Some(spec));
}
if let Some(declared) = declared
&& let Some(spec) = SortSpec::parse(declared).context("read the table's declared order")?
{
return Ok(Some(spec));
}
SortSpec::parse(configured).context("read the configured sort order")
}