Files
seaweedfs/seaweed-volume/src/storage/volume.rs
T
10b0f2b8ad volume server: refuse the rest of a grouped run after a durable index failure (#11576)
* volume server: refuse the rest of a grouped run after a durable index failure

A durable write whose needle-map put fails stops the volume taking writes
(#10825): sent on its own, the next write then fails read only before it
appends. The grouped run from #11543 appends and syncs every entry before
publishing any, then kept publishing the entries after the failed one and
acked them once the shared .idx sync went through. When the failed put
tore its .idx row, the rows appended after it land off alignment, so the
next load parses them as garbage and the acked writes are gone.

Once a durable entry fails to publish, refuse every later entry of the run
with ReadOnly, as the per-needle path does. The entries before it stay
acked; their rows go down with the run's one .idx sync. The refused
records stay on the .dat unindexed, as the failed one does on its own.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: refuse a grouped entry staged as a cookie mismatch too

After a durable entry in a grouped run fails to index, the entries
after it are refused as they would be on their own. On its own an entry
meets check_writable before its cookie check, so one staged as a cookie
mismatch now gets the refusal too, instead of keeping its staging error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: trim a torn .idx row back so the next stays aligned

A failed write_index_entry can leave half a row in the .idx. With the
writer appending at the tail, every row written after it lands off
alignment and the next load parses them as garbage, so a write acked
behind a torn row does not come back. Trim the file back to
idx_file_offset on a failed append, in both needle maps, and cover it
with a test that writes past a torn row and reloads.

* volume server: refuse queued Go writes once a durable index update fails

processBatch kept writing after a failed nm.Put, and the single-write
path checked IsReadOnly only outside the volume lock. A durable write
whose index update fails now marks the volume noWriteOrDelete, and each
queued request is checked before it appends, so the ones after a failed
durable entry are refused the way a lone write is. Deletes get the same
noWriteOrDelete refusal a lone delete gets.

* volume server: refuse appends while a torn .idx row cannot be trimmed

When trimming back a half-written .idx row itself fails, the next append
would land after the torn bytes and every later row would parse off
alignment on load. Latch the map as torn and refuse appends until the
trim succeeds, on both CompactNeedleMap and RedbNeedleMap; the same
latch covers an orphan row that could not be trimmed after a failed
redb commit.

The .idx writer is now opened with write+append access so truncate_to
(set_len) works on Windows, where an append-only handle cannot trim.

* volume server: write .idx rows at idx_file_offset, not via append mode

Rust's OpenOptions on Windows strips FILE_WRITE_DATA whenever append is
set so the handle stays strictly append-only, which makes set_len fail -
the torn-row trim could never succeed there. Open the .idx writer with
plain write access and seek to idx_file_offset before each row, the same
positioned-write model the Go server uses.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 21:04:30 +08:00

10076 lines
368 KiB
Rust

//! Volume: the core storage unit — a .dat file + .idx index.
//!
//! Each volume contains many needles (files). It manages:
//! - Reading/writing/deleting needles from the .dat file
//! - Maintaining the in-memory NeedleMap (NeedleId → Offset+Size)
//! - SuperBlock at offset 0 of the .dat file
//! - Metrics (file count, content size, deleted count)
//!
//! Matches Go's storage/volume.go, volume_loading.go, volume_read.go,
//! volume_write.go, volume_super_block.go.
use std::collections::HashSet;
use std::fs::{self, File, OpenOptions};
use std::io::{self, Read, Seek, SeekFrom, Write};
use std::ops::ControlFlow;
use std::path::Path;
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{Condvar, Mutex};
use std::time::{SystemTime, UNIX_EPOCH};
use tracing::{debug, error, info, warn};
use crate::storage::idx;
use crate::storage::io::read_exact_at;
use crate::storage::io_error::{IoErrorTracker, StreakMark};
use crate::storage::needle::needle::{self, Needle, NeedleError, get_actual_size};
use crate::storage::needle_map::sorted_file::SortedFileNeedleMap;
use crate::storage::needle_map::{CompactNeedleMap, NeedleMap, NeedleMapKind, RedbNeedleMap};
use crate::storage::super_block::{ReplicaPlacement, SUPER_BLOCK_SIZE, SuperBlock};
use crate::storage::types::*;
use crate::storage::volume_open::open_volume_file;
// ============================================================================
// Errors
// ============================================================================
#[derive(Debug, thiserror::Error)]
pub enum VolumeError {
#[error("not found")]
NotFound,
#[error("volume id {0} is not found")]
VolumeNotFound(VolumeId),
#[error("already deleted")]
Deleted,
#[error("needle size mismatch")]
SizeMismatch,
#[error("unsupported version: {0}")]
UnsupportedVersion(u8),
#[error("cookie mismatch: {0:#x}")]
CookieMismatch(u32),
#[error("volume not empty")]
NotEmpty,
#[error("volume already exists")]
AlreadyExists,
#[error("volume {0} is read only")]
ReadOnly(VolumeId),
#[error("volume is unavailable: {0}")]
Unavailable(String),
#[error("volume size limit exceeded: current {current}, limit {limit}")]
SizeLimitExceeded { current: u64, limit: u64 },
#[error("not enough free space: required {required}, free {free}")]
InsufficientSpace {
vid: VolumeId,
required: u64,
free: u64,
},
#[error("volume not initialized")]
NotInitialized,
#[error("needle error: {0}")]
Needle(#[from] NeedleError),
#[error("super block error: {0}")]
SuperBlock(#[from] crate::storage::super_block::SuperBlockError),
#[error("IO error: {0}")]
Io(#[from] io::Error),
#[error(transparent)]
Tier(#[from] crate::remote_storage::s3_tier::TierError),
}
/// Returns true when a needle read failed because the on-disk bytes are
/// unreadable in a permanent way (offset past EOF, header corruption, CRC
/// mismatch, malformed v2/v3/v4 fields). Vacuum can safely drop such entries
/// during compaction. Anything else (real disk EIO, ERROR_CRC on Windows,
/// network timeouts, EROFS, etc.) is transient or environmental and must
/// abort the compaction so an operator notices.
fn is_skippable_needle_read_error(e: &VolumeError) -> bool {
match e {
VolumeError::Io(io_err) => io_err.kind() == io::ErrorKind::UnexpectedEof,
VolumeError::Needle(NeedleError::SizeMismatch { .. }) => true,
VolumeError::Needle(NeedleError::CrcMismatch { .. }) => true,
VolumeError::Needle(NeedleError::IndexOutOfRange(_)) => true,
VolumeError::Needle(NeedleError::TailTooShort) => true,
VolumeError::SizeMismatch => true,
_ => false,
}
}
/// Reports whether the compacted .dat is short of the live bytes the
/// pre-compaction index snapshot expected.
fn exceeds_expected_compacted_size(expected_live_bytes: u64, dst_dat_size: u64) -> bool {
expected_live_bytes > dst_dat_size
}
/// Read and parse the needle at `offset` through `read_at`.
fn read_needle_with(
read_at: impl Fn(&mut [u8], u64) -> Result<(), VolumeError>,
n: &mut Needle,
offset: i64,
size: Size,
version: Version,
) -> Result<(), VolumeError> {
match parse_needle_at(&read_at, n, offset, size, version) {
Ok(()) => Ok(()),
#[cfg(not(feature = "5bytes"))]
Err(VolumeError::Needle(NeedleError::SizeMismatch { offset: o, .. }))
if o < MAX_POSSIBLE_VOLUME_SIZE as i64 =>
{
// Double-read: in 4-byte offset mode, the actual data may be
// beyond 32GB due to offset wrapping. Retry at offset + 32GB.
parse_needle_at(
&read_at,
n,
offset + MAX_POSSIBLE_VOLUME_SIZE as i64,
size,
version,
)
}
Err(e) => Err(e),
}
}
fn parse_needle_at(
read_at: &impl Fn(&mut [u8], u64) -> Result<(), VolumeError>,
n: &mut Needle,
offset: i64,
size: Size,
version: Version,
) -> Result<(), VolumeError> {
// Storage guard: negativity-only (Go parity — storage allocates what the
// index says). Size(0) and >1GiB map sizes must still read; only
// negative wraps/panics. Transport cap lives in RPC handlers only.
if size.0 < 0 {
return Err(VolumeError::Io(std::io::Error::new(
std::io::ErrorKind::InvalidData,
format!("invalid needle size {}", size.0),
)));
}
let actual_size = get_actual_size(size, version);
let mut buf = vec![0u8; actual_size as usize];
#[cfg(test)]
needle_read_hook::fire(n.id, buf.len());
read_at(&mut buf, offset as u64)?;
n.read_bytes(&buf, offset, size, version)?;
Ok(())
}
// ============================================================================
// VolumeInfo (.vif persistence)
// ============================================================================
/// Legacy simple VolumeInfo for backward compat with old .vif files.
#[derive(serde::Serialize, serde::Deserialize)]
struct VolumeInfo {
read_only: bool,
}
pub use crate::pb::volume_server_pb::RemoteFile as PbRemoteFile;
/// Protobuf VolumeInfo type alias.
pub use crate::pb::volume_server_pb::VolumeInfo as PbVolumeInfo;
/// Helper module for deserializing protojson uint64 fields that may be strings or numbers.
mod string_or_u64 {
use serde::{self, Deserialize, Deserializer, Serializer};
pub fn serialize<S>(value: &u64, serializer: S) -> Result<S::Ok, S::Error>
where
S: Serializer,
{
// Emit as string to match Go's protojson format for uint64
serializer.serialize_str(&value.to_string())
}
pub fn deserialize<'de, D>(deserializer: D) -> Result<u64, D::Error>
where
D: Deserializer<'de>,
{
#[derive(Deserialize)]
#[serde(untagged)]
enum StringOrNum {
Str(String),
Num(u64),
}
match StringOrNum::deserialize(deserializer)? {
StringOrNum::Str(s) => s.parse::<u64>().map_err(serde::de::Error::custom),
StringOrNum::Num(n) => Ok(n),
}
}
}
mod string_or_i64 {
use serde::{self, Deserialize, Deserializer, Serializer};
pub fn serialize<S>(value: &i64, serializer: S) -> Result<S::Ok, S::Error>
where
S: Serializer,
{
serializer.serialize_str(&value.to_string())
}
pub fn deserialize<'de, D>(deserializer: D) -> Result<i64, D::Error>
where
D: Deserializer<'de>,
{
#[derive(Deserialize)]
#[serde(untagged)]
enum StringOrNum {
Str(String),
Num(i64),
}
match StringOrNum::deserialize(deserializer)? {
StringOrNum::Str(s) => s.parse::<i64>().map_err(serde::de::Error::custom),
StringOrNum::Num(n) => Ok(n),
}
}
}
/// Serde-compatible representation of RemoteFile for .vif JSON serialization.
/// Field names use snake_case to match Go's protobuf JSON output (jsonpb).
#[derive(serde::Serialize, serde::Deserialize, Default, Clone)]
pub struct VifRemoteFile {
#[serde(default, rename = "backendType")]
pub backend_type: String,
#[serde(default, rename = "backendId")]
pub backend_id: String,
#[serde(default)]
pub key: String,
#[serde(default, with = "string_or_u64")]
pub offset: u64,
#[serde(default, rename = "fileSize", with = "string_or_u64")]
pub file_size: u64,
#[serde(default, rename = "modifiedTime", with = "string_or_u64")]
pub modified_time: u64,
#[serde(default)]
pub extension: String,
}
#[derive(serde::Serialize, serde::Deserialize, Default, Clone)]
pub struct VifEcShardConfig {
#[serde(default, rename = "dataShards")]
pub data_shards: u32,
#[serde(default, rename = "parityShards")]
pub parity_shards: u32,
/// Encode time (unix nanos). Shards and .ecx of one encode run share it,
/// so a read served from a different run's shard is rejected.
#[serde(default, rename = "encodeTsNs", with = "string_or_i64")]
pub encode_ts_ns: i64,
/// Uniform block layout: each shard is a single contiguous block of this
/// many bytes. 0 = legacy 1GiB/1MiB two-tier layout.
#[serde(default, rename = "blockSize", with = "string_or_i64")]
pub block_size: i64,
}
/// Serde-compatible representation of OldVersionVolumeInfo for legacy .vif JSON deserialization.
/// Matches Go's protobuf OldVersionVolumeInfo where `DestroyTime` maps to `expire_at_sec`.
#[derive(serde::Deserialize, Default)]
struct OldVersionVifVolumeInfo {
#[serde(default)]
pub files: Vec<VifRemoteFile>,
#[serde(default)]
pub version: u32,
#[serde(default)]
pub replication: String,
#[serde(default, alias = "bytesOffset", alias = "BytesOffset")]
pub bytes_offset: u32,
#[serde(
default,
alias = "datFileSize",
alias = "dat_file_size",
with = "string_or_i64"
)]
pub dat_file_size: i64,
#[serde(
default,
alias = "destroyTime",
alias = "DestroyTime",
with = "string_or_u64"
)]
pub destroy_time: u64,
#[serde(default, alias = "readOnly", alias = "read_only")]
pub read_only: bool,
}
impl OldVersionVifVolumeInfo {
/// Convert to the standard VifVolumeInfo, mapping destroy_time -> expire_at_sec.
fn into_vif(self) -> VifVolumeInfo {
VifVolumeInfo {
files: self.files,
version: self.version,
replication: self.replication,
bytes_offset: self.bytes_offset,
dat_file_size: self.dat_file_size,
expire_at_sec: self.destroy_time,
read_only: self.read_only,
ec_shard_config: None,
read_only_can_delete: false,
}
}
}
/// Serde-compatible representation of VolumeInfo for .vif JSON serialization.
/// Matches Go's protobuf JSON format (jsonpb with EmitUnpopulated=true).
#[derive(serde::Serialize, serde::Deserialize, Default, Clone)]
pub struct VifVolumeInfo {
#[serde(default)]
pub files: Vec<VifRemoteFile>,
#[serde(default)]
pub version: u32,
#[serde(default)]
pub replication: String,
#[serde(default, rename = "bytesOffset")]
pub bytes_offset: u32,
#[serde(default, rename = "datFileSize", with = "string_or_i64")]
pub dat_file_size: i64,
#[serde(default, rename = "expireAtSec", with = "string_or_u64")]
pub expire_at_sec: u64,
#[serde(default, rename = "readOnly")]
pub read_only: bool,
#[serde(
default,
rename = "ecShardConfig",
skip_serializing_if = "Option::is_none"
)]
pub ec_shard_config: Option<VifEcShardConfig>,
#[serde(default, rename = "readOnlyCanDelete")]
pub read_only_can_delete: bool,
}
impl VifVolumeInfo {
/// Convert from protobuf VolumeInfo to the serde-compatible struct.
pub fn from_pb(pb: &PbVolumeInfo) -> Self {
Self {
files: pb
.files
.iter()
.map(|f| VifRemoteFile {
backend_type: f.backend_type.clone(),
backend_id: f.backend_id.clone(),
key: f.key.clone(),
offset: f.offset,
file_size: f.file_size,
modified_time: f.modified_time,
extension: f.extension.clone(),
})
.collect(),
version: pb.version,
replication: pb.replication.clone(),
bytes_offset: pb.bytes_offset,
dat_file_size: pb.dat_file_size,
expire_at_sec: pb.expire_at_sec,
read_only: pb.read_only,
ec_shard_config: pb.ec_shard_config.as_ref().map(|c| VifEcShardConfig {
data_shards: c.data_shards,
parity_shards: c.parity_shards,
encode_ts_ns: c.encode_ts_ns,
block_size: c.block_size,
}),
read_only_can_delete: pb.read_only_can_delete,
}
}
/// Convert to protobuf VolumeInfo.
pub fn to_pb(&self) -> PbVolumeInfo {
PbVolumeInfo {
files: self
.files
.iter()
.map(|f| PbRemoteFile {
backend_type: f.backend_type.clone(),
backend_id: f.backend_id.clone(),
key: f.key.clone(),
offset: f.offset,
file_size: f.file_size,
modified_time: f.modified_time,
extension: f.extension.clone(),
})
.collect(),
version: self.version,
replication: self.replication.clone(),
bytes_offset: self.bytes_offset,
dat_file_size: self.dat_file_size,
expire_at_sec: self.expire_at_sec,
read_only: self.read_only,
ec_shard_config: self.ec_shard_config.as_ref().map(|c| {
crate::pb::volume_server_pb::EcShardConfig {
data_shards: c.data_shards,
parity_shards: c.parity_shards,
encode_ts_ns: c.encode_ts_ns,
block_size: c.block_size,
}
}),
read_only_can_delete: self.read_only_can_delete,
}
}
}
// ============================================================================
// Streaming read support
// ============================================================================
#[derive(Default)]
struct DataFileAccessState {
readers: usize,
writer_active: bool,
}
#[derive(Default)]
pub struct DataFileAccessControl {
state: Mutex<DataFileAccessState>,
condvar: Condvar,
}
pub struct DataFileReadLease {
control: Arc<DataFileAccessControl>,
}
pub struct DataFileWriteLease {
control: Arc<DataFileAccessControl>,
}
impl DataFileAccessControl {
pub fn read_lock(self: &Arc<Self>) -> DataFileReadLease {
let mut state = self.state.lock().unwrap();
while state.writer_active {
state = self.condvar.wait(state).unwrap();
}
state.readers += 1;
drop(state);
DataFileReadLease {
control: self.clone(),
}
}
pub fn write_lock(self: &Arc<Self>) -> DataFileWriteLease {
let mut state = self.state.lock().unwrap();
while state.writer_active || state.readers > 0 {
state = self.condvar.wait(state).unwrap();
}
state.writer_active = true;
drop(state);
DataFileWriteLease {
control: self.clone(),
}
}
}
impl Drop for DataFileReadLease {
fn drop(&mut self) {
let mut state = self.control.state.lock().unwrap();
state.readers -= 1;
if state.readers == 0 {
self.control.condvar.notify_all();
}
}
}
impl Drop for DataFileWriteLease {
fn drop(&mut self) {
let mut state = self.control.state.lock().unwrap();
state.writer_active = false;
self.control.condvar.notify_all();
}
}
/// Test seam: a hook sees (and may park) every data read of its needle.
#[cfg(test)]
pub(crate) mod needle_read_hook {
use super::NeedleId;
use std::sync::{Arc, Mutex};
type Hook = Arc<dyn Fn(usize) + Send + Sync>;
static HOOKS: Mutex<Vec<(NeedleId, Hook)>> = Mutex::new(Vec::new());
pub(crate) struct Registration(NeedleId);
impl Drop for Registration {
fn drop(&mut self) {
HOOKS.lock().unwrap().retain(|(id, _)| *id != self.0);
}
}
/// Needle ids are the key, so tests running in parallel must use distinct ones.
pub(crate) fn register(
id: NeedleId,
hook: impl Fn(usize) + Send + Sync + 'static,
) -> Registration {
HOOKS.lock().unwrap().push((id, Arc::new(hook)));
Registration(id)
}
pub(crate) fn fire(id: NeedleId, len: usize) {
let hook = HOOKS
.lock()
.unwrap()
.iter()
.find(|(hook_id, _)| *hook_id == id)
.map(|(_, hook)| hook.clone());
if let Some(hook) = hook {
hook(len);
}
}
}
/// Information needed to stream needle data directly from the dat file
/// without loading the entire payload into memory.
pub(crate) enum NeedleStreamSource {
Local(File),
Remote(RemoteDatFile),
}
impl NeedleStreamSource {
pub(crate) fn read_exact_at(&self, buf: &mut [u8], offset: u64) -> io::Result<()> {
match self {
NeedleStreamSource::Local(file) => read_exact_at(file, buf, offset),
NeedleStreamSource::Remote(remote) => remote.read_exact_at(buf, offset),
}
}
}
/// A `.dat` scan that runs with the store lock released, built by
/// `Volume::dat_scan_plan` while the caller holds a store guard.
///
/// The start offset, the handle and the end bound are captured together
/// while the guard excludes every writer. The handle pins the inode the
/// offset was resolved against: a vacuum commit replaces `.dat` by rename,
/// and unmount or destroy close and unlink it, none of which rewrites the
/// pinned bytes. `.dat` is otherwise append-only and a failed append
/// truncates only its own bytes, so `[from, end)` holds complete records
/// that nothing rewrites while the scan runs. A short read below `end` can
/// therefore only mean the inode was truncated under the plan (an unmount
/// followed by a `VolumeCopy` of the same id reopens `.dat` with
/// `truncate(true)`), and the scan fails rather than ending as if the
/// snapshot had been read. The fields are private so a plan can only come
/// from `dat_scan_plan`.
pub(crate) struct DatScanPlan {
source: NeedleStreamSource,
version: Version,
from: u64,
end: u64,
/// The volume instance and revision the plan was taken from: after a
/// vacuum or a re-create, the needle map's offsets name another file.
owner: Arc<DataFileAccessControl>,
compaction_revision: u16,
}
/// One record visited by `DatScanPlan::scan`: the raw on-disk bytes and the
/// parsed append timestamp.
pub(crate) struct RawNeedle<'a> {
pub header: &'a [u8],
pub body: &'a [u8],
pub append_at_ns: u64,
}
/// One record visited by `DatScanPlan::scan_records`, not yet parsed.
pub(crate) struct DatRecord<'a> {
pub offset: u64,
pub id: NeedleId,
pub size: Size,
/// Header and body.
pub bytes: &'a [u8],
version: Version,
}
impl DatRecord<'_> {
pub(crate) fn parse(&self) -> Result<Needle, VolumeError> {
let mut n = Needle::default();
n.read_bytes(self.bytes, self.offset as i64, self.size, self.version)?;
Ok(n)
}
}
impl DatScanPlan {
/// Visit the records in `[from, end)` in file order, the way Go's
/// `ScanVolumeFileFrom` feeds a `VolumeFileScanner`: only the record being
/// visited is in memory. Reads are positional and never touch the
/// `Volume`. The pass ends `Ok` at the end of the data, at a record that
/// does not fit before `end`, or when `visit` breaks. A corrupt header
/// fails the pass: ending early would let a truncated tail look complete.
/// It also fails on any read error, including a short read below `end`:
/// every byte below `end` existed when the plan was taken, so a short read
/// means the file was truncated under the plan and the pass is incomplete.
pub(crate) fn scan(
&self,
mut visit: impl FnMut(RawNeedle<'_>) -> ControlFlow<()>,
) -> Result<(), VolumeError> {
self.scan_records(|record| {
let append_at_ns = record.parse()?.append_at_ns;
Ok(visit(RawNeedle {
header: &record.bytes[..NEEDLE_HEADER_SIZE],
body: &record.bytes[NEEDLE_HEADER_SIZE..],
append_at_ns,
}))
})
}
/// `scan` without the parse: a visitor that skips a record never fails on
/// a body it did not need.
pub(crate) fn scan_records(
&self,
mut visit: impl FnMut(DatRecord<'_>) -> Result<ControlFlow<()>, VolumeError>,
) -> Result<(), VolumeError> {
let mut offset = self.from;
while offset + NEEDLE_HEADER_SIZE as u64 <= self.end {
let mut header = [0u8; NEEDLE_HEADER_SIZE];
self.source
.read_exact_at(&mut header, offset)
.map_err(|e| self.read_error(e, offset))?;
let (_cookie, id, size) = Needle::parse_header(&header);
if size.0 == 0 && id.is_empty() {
break;
}
// A negative size is a corrupt header: the record length it
// implies cannot advance the scan past it, and stopping quietly
// would let a truncated tail look complete. Go's scan fails here.
if size.0 < 0 {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"corrupt needle header at offset {offset}: size {}, record length {}",
size.0,
get_actual_size(size, self.version)
),
)));
}
// Nothing past `end` was a complete record when the plan was
// taken. The size is checked against the bytes left, and the
// whole record after it. Both run before allocating, so a corrupt
// size cannot size the buffer. (The size check alone does not
// keep a corrupt size out of the padding arithmetic on a volume
// with more than 2 GiB left, which is why `padding_length` sums
// in i64.)
if size.0 as u64 > self.end - offset - NEEDLE_HEADER_SIZE as u64 {
break;
}
let body_length = needle::needle_body_length(size, self.version) as u64;
let total_size = NEEDLE_HEADER_SIZE as u64 + body_length;
if offset + total_size > self.end {
break;
}
// Match Go's ScanVolumeFileFrom: visit ALL needles including deleted ones.
// This is critical for incremental copy where tombstones must be propagated.
let mut record = vec![0u8; total_size as usize];
record[..NEEDLE_HEADER_SIZE].copy_from_slice(&header);
self.source
.read_exact_at(
&mut record[NEEDLE_HEADER_SIZE..],
offset + NEEDLE_HEADER_SIZE as u64,
)
.map_err(|e| self.read_error(e, offset + NEEDLE_HEADER_SIZE as u64))?;
let record = DatRecord {
offset,
id,
size,
bytes: &record,
version: self.version,
};
if visit(record)?.is_break() {
break;
}
offset += total_size;
}
Ok(())
}
/// Name the snapshot in a read failure. A short read is the truncation
/// case described on the type; every other error is passed through with
/// the offset.
fn read_error(&self, e: io::Error, offset: u64) -> VolumeError {
let what = if e.kind() == io::ErrorKind::UnexpectedEof {
"short read below the snapshot end, the dat file was truncated under the scan"
} else {
"read failed"
};
VolumeError::Io(io::Error::new(
e.kind(),
format!(
"{} (offset {}, snapshot end {}): {}",
what, offset, self.end, e
),
))
}
}
/// Why a volume refuses all I/O, shared with its in-flight needle streams so
/// they see a mark made after they left the store lock. A leaf lock.
#[derive(Debug, Default)]
pub(crate) struct IoUnavailable(Mutex<Option<String>>);
impl IoUnavailable {
fn set(&self, reason: String) {
*self.0.lock().unwrap() = Some(reason);
}
fn is_set(&self) -> bool {
self.0.lock().unwrap().is_some()
}
pub(crate) fn error(&self) -> Option<VolumeError> {
self.0
.lock()
.unwrap()
.as_ref()
.map(|reason| VolumeError::Unavailable(reason.clone()))
}
}
/// A needle read resolved under a store guard and run after it is released;
/// the handle pins the inode, as for `DatScanPlan`. It takes no data-file
/// lease: writers wait for one while holding the store write lock.
pub(crate) struct NeedleReadPlan {
source: NeedleStreamSource,
offset: i64,
size: Size,
version: Version,
needle_id: NeedleId,
data_file_access_control: Arc<DataFileAccessControl>,
io_errors: Arc<IoErrorTracker>,
io_unavailable: Arc<IoUnavailable>,
}
impl NeedleReadPlan {
/// The needle's size in the index.
pub(crate) fn size(&self) -> Size {
self.size
}
/// Read the whole needle and verify its checksum.
pub(crate) fn read_full(&mut self, n: &mut Needle) -> Result<(), VolumeError> {
let (source, size, version) = (&self.source, self.size, self.version);
let read_at = |buf: &mut [u8], at: u64| Ok(source.read_exact_at(buf, at)?);
let result = with_offset_retry(&mut self.offset, |off| {
parse_needle_at(&read_at, n, off, size, version)
});
match &result {
Ok(()) => self.io_errors.check_read_write_error(None),
Err(VolumeError::Io(e)) => self.io_errors.check_read_write_error(Some(e)),
Err(_) => {}
}
result?;
if needle_expired(n) {
return Err(VolumeError::NotFound);
}
Ok(())
}
/// Read only the needle's metadata, skipping the payload.
pub(crate) fn read_meta(&mut self, n: &mut Needle) -> Result<(), VolumeError> {
let (source, size, version) = (&self.source, self.size, self.version);
let read_at = |buf: &mut [u8], at: u64| Ok(source.read_exact_at(buf, at)?);
with_offset_retry(&mut self.offset, |off| {
if version == VERSION_1 {
// A V1 body is all data, so its "meta" is the whole record.
let mut buf = vec![0u8; get_actual_size(size, version) as usize];
#[cfg(test)]
needle_read_hook::fire(n.id, buf.len());
read_at(&mut buf, off as u64)?;
Ok(n.read_bytes_meta_only(&buf, off, size, version)?)
} else {
Volume::read_needle_meta_blob_and_parse(read_at, n, off, size, version)
}
})?;
if needle_expired(n) {
return Err(VolumeError::NotFound);
}
Ok(())
}
/// For a reply served from the data file, after `read_meta` filled `n`.
pub(crate) fn into_stream_info(self, n: &Needle) -> NeedleStreamInfo {
// V1 data starts right after the header, V2/V3 after the DataSize field.
let data_file_offset = if self.version == VERSION_1 {
self.offset as u64 + NEEDLE_HEADER_SIZE as u64
} else {
self.offset as u64 + NEEDLE_HEADER_SIZE as u64 + DATA_SIZE_SIZE as u64
};
NeedleStreamInfo {
source: self.source,
data_file_offset,
data_size: n.data_size,
data_file_access_control: self.data_file_access_control,
io_unavailable: self.io_unavailable,
needle_id: self.needle_id,
checksum: n.checksum.0,
}
}
}
/// Run `read` at `offset`, retrying past the 32 GiB wrap of 4-byte offsets
/// on a size mismatch, and leave `offset` where the needle parsed.
fn with_offset_retry(
offset: &mut i64,
mut read: impl FnMut(i64) -> Result<(), VolumeError>,
) -> Result<(), VolumeError> {
match read(*offset) {
#[cfg(not(feature = "5bytes"))]
Err(VolumeError::Needle(NeedleError::SizeMismatch { offset: o, .. }))
if o < MAX_POSSIBLE_VOLUME_SIZE as i64 =>
{
*offset += MAX_POSSIBLE_VOLUME_SIZE as i64;
read(*offset)
}
result => result,
}
}
fn needle_expired(n: &Needle) -> bool {
let Some(ttl) = n.ttl.as_ref().filter(|_| n.has_ttl()) else {
return false;
};
let ttl_minutes = ttl.minutes();
if ttl_minutes == 0 || !n.has_last_modified_date() {
return false;
}
let expire_at_ns = n.append_at_ns + (ttl_minutes as u64) * 60 * 1_000_000_000;
let now_ns = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_nanos() as u64;
now_ns >= expire_at_ns
}
/// Holds `Volume::is_compacting` set; dropping it clears the flag.
struct CompactionClaim(Arc<AtomicBool>);
impl CompactionClaim {
fn try_claim(flag: &Arc<AtomicBool>) -> Option<Self> {
flag.compare_exchange(false, true, Ordering::AcqRel, Ordering::Acquire)
.ok()
.map(|_| CompactionClaim(flag.clone()))
}
}
impl Drop for CompactionClaim {
fn drop(&mut self) {
self.0.store(false, Ordering::Release);
}
}
/// A compaction copy that runs off the store lock: `.dat`/`.idx` are
/// append-only below `idx_size`, and `makeup_diff` replays what lands after.
pub(crate) struct CompactionJob {
id: VolumeId,
src_dat: NeedleStreamSource,
src_idx: File,
idx_size: u64,
cpd_path: String,
cpx_path: String,
version: Version,
super_block: SuperBlock,
io_errors: Arc<IoErrorTracker>,
_claim: CompactionClaim,
}
impl CompactionJob {
pub(crate) fn run<F>(self, progress_fn: F) -> Result<(), VolumeError>
where
F: Fn(i64) -> bool,
{
let cpd_path = &self.cpd_path;
let cpx_path = &self.cpx_path;
let version = self.version;
// Write new super block with incremented compaction revision
let mut new_sb = self.super_block.clone();
new_sb.compaction_revision += 1;
let sb_bytes = new_sb.to_bytes();
let mut dst = OpenOptions::new()
.read(true)
.write(true)
.create(true)
.truncate(true)
.open(cpd_path)?;
dst.write_all(&sb_bytes)?;
let mut new_offset = sb_bytes.len() as i64;
// Build new index in memory
let mut new_nm = CompactNeedleMap::new();
let now = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_secs();
let entries = self.live_entries()?;
let mut skipped_needles: u64 = 0;
let mut skipped_data_bytes: u64 = 0;
let mut expected_live_bytes: u64 = 0;
for (id, offset, size) in entries {
// Progress callback
if !progress_fn(offset.to_actual_offset()) {
// Interrupted
let _ = fs::remove_file(cpd_path);
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Interrupted,
"compaction interrupted",
)));
}
// Read needle from source
let mut n = Needle {
id,
..Needle::default()
};
let read = read_needle_with(
|buf, at| self.src_dat.read_exact_at(buf, at).map_err(VolumeError::Io),
&mut n,
offset.to_actual_offset(),
size,
version,
);
if let Err(e) = read {
// Record EIO for health monitoring (parity with Go's checkReadWriteError).
if let VolumeError::Io(ref io_err) = e {
self.io_errors.check_read_write_error(Some(io_err));
}
// Only drop the entry when the failure is one of the well-
// known permanent-corruption shapes. A transient disk fault,
// a tiered-read timeout, or a Windows hardware error (which
// surfaces as a generic Io rather than UnexpectedEof) must
// abort so an operator notices, rather than silently
// compacting away data that might come back on retry.
if !is_skippable_needle_read_error(&e) {
return Err(VolumeError::Io(io::Error::other(format!(
"cannot hydrate needle from file: {}",
e
))));
}
skipped_needles += 1;
if size.is_valid() {
skipped_data_bytes += size.0 as u64;
}
warn!(
volume_id = self.id.0,
key = id.0,
offset = offset.to_actual_offset(),
size = size.0,
error = %e,
"vacuum: dropping unreadable needle"
);
continue;
}
// Skip TTL-expired needles using the volume's TTL (matches Go's volume_vacuum.go)
if n.has_ttl() {
let ttl_minutes = self.super_block.ttl.minutes();
if ttl_minutes > 0 && n.last_modified > 0 {
let expire_at = n.last_modified + (ttl_minutes as u64) * 60;
if now >= expire_at {
continue;
}
}
}
// Tally the live bytes from the frozen snapshot this loop copied
// from, not the live needle map. Unreadable needles return before
// this point, so no further skipped-byte adjustment is needed.
expected_live_bytes += size.0 as u64;
// Write needle to destination
let bytes = n.write_bytes(version);
dst.write_all(&bytes)?;
// Update new index
new_nm.put(id, Offset::from_actual_offset(new_offset), n.size)?;
new_offset += bytes.len() as i64;
}
if skipped_needles > 0 {
warn!(
volume_id = self.id.0,
skipped_needles,
skipped_data_bytes,
"vacuum: dropped unreadable index entries during compaction"
);
}
dst.sync_all()?;
if self.super_block.ttl.is_empty() {
let dst_dat_size = dst.metadata()?.len();
if exceeds_expected_compacted_size(expected_live_bytes, dst_dat_size) {
let _ = fs::remove_file(cpd_path);
let _ = fs::remove_file(cpx_path);
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::UnexpectedEof,
format!(
"volume {} unexpected new data size: {} does not match expected live content size {} from the pre-compaction snapshot",
self.id.0, dst_dat_size, expected_live_bytes
),
)));
}
}
// Save new index
new_nm.save_to_idx(cpx_path)?;
Ok(())
}
/// Replay `.idx` up to the recorded size, as Go's `LoadFromIdx` does, and
/// return the live entries in `.dat` order. A short read fails: treating
/// it as the whole index would compact away every needle past it.
fn live_entries(&self) -> Result<Vec<(NeedleId, Offset, Size)>, VolumeError> {
let rows = self.idx_size / NEEDLE_MAP_ENTRY_SIZE as u64;
let mut snapshot = CompactNeedleMap::new();
let mut seen = 0u64;
// Rows past the recorded size are makeup_diff's; do not read them.
let mut reader = io::BufReader::new((&self.src_idx).take(self.idx_size));
let mut buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
while seen < rows {
match reader.read_exact(&mut buf) {
Ok(()) => {}
Err(e) if e.kind() == io::ErrorKind::UnexpectedEof => break,
Err(e) => return Err(e.into()),
}
let (key, offset, size) = idx_entry_from_bytes(&buf);
if offset.is_zero() || size.is_deleted() {
snapshot.delete(key, offset)?;
} else {
snapshot.put(key, offset, size)?;
}
seen += 1;
}
if seen < rows {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::UnexpectedEof,
format!(
"volume {} index has {} entries, expected {}",
self.id.0, seen, rows
),
)));
}
let mut entries: Vec<(NeedleId, Offset, Size)> = Vec::new();
let _ = snapshot.ascending_visit(|id, nv| {
if !nv.offset.is_zero() && !nv.size.is_deleted() {
entries.push((id, nv.offset, nv.size));
}
Ok::<(), std::convert::Infallible>(())
});
entries.sort_by_key(|(_, offset, _)| *offset);
Ok(entries)
}
}
pub struct NeedleStreamInfo {
/// Stream source for the dat file, local or remote.
pub(crate) source: NeedleStreamSource,
/// Absolute byte offset within the dat file where needle data starts.
pub data_file_offset: u64,
/// Size of the data payload in bytes.
pub data_size: u32,
/// Per-volume file access lock used to match Go's slow-read behavior.
pub data_file_access_control: Arc<DataFileAccessControl>,
/// Checked before each chunk: the volume can become unavailable mid-stream.
pub(crate) io_unavailable: Arc<IoUnavailable>,
pub needle_id: NeedleId,
/// Checksum stored in the needle tail, verified once the last chunk has
/// been read — before that frame is emitted.
pub checksum: u32,
}
#[derive(Clone)]
pub(crate) struct RemoteDatFile {
backend: Arc<crate::remote_storage::s3_tier::S3TierBackend>,
key: String,
file_size: u64,
modified_time: u64,
}
impl RemoteDatFile {
/// Read up to `size` bytes at `offset`, bounded by the remote file size.
/// Mirrors `Volume::read_dat_slice` for a remote-only backend, but holds no
/// volume/store lock so callers can stream off the store lock entirely.
pub(crate) fn read_slice(&self, offset: u64, size: usize) -> io::Result<Vec<u8>> {
if size == 0 || offset >= self.file_size {
return Ok(Vec::new());
}
let read_len = std::cmp::min(size as u64, self.file_size - offset) as usize;
let mut buf = vec![0u8; read_len];
self.read_exact_at(&mut buf, offset)?;
Ok(buf)
}
pub(crate) fn read_exact_at(&self, buf: &mut [u8], offset: u64) -> io::Result<()> {
let data = self
.backend
.read_range_blocking(&self.key, offset, buf.len())
.map_err(io::Error::other)?;
if data.len() != buf.len() {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
format!(
"remote read short read at offset {}: got {}, expected {}",
offset,
data.len(),
buf.len()
),
));
}
buf.copy_from_slice(&data);
Ok(())
}
}
// ============================================================================
// Volume
// ============================================================================
pub struct Volume {
pub id: VolumeId,
dir: String,
dir_idx: String,
pub collection: String,
dat_file: Option<File>,
remote_dat_file: Option<RemoteDatFile>,
nm: Option<NeedleMap>,
/// Make the next .dat or .idx flush fail, so tests can exercise the durable
/// write's failure paths without a real disk fault.
#[cfg(test)]
fail_fsync_for_test: bool,
#[cfg(test)]
fail_idx_sync_for_test: bool,
#[cfg(test)]
fail_truncate_for_test: bool,
/// Needle ids whose .dat append fails with a media error.
#[cfg(test)]
fail_append_for_test: HashSet<NeedleId>,
#[cfg(test)]
dat_syncs_for_test: std::sync::atomic::AtomicUsize,
#[cfg(test)]
idx_syncs_for_test: std::sync::atomic::AtomicUsize,
needle_map_kind: NeedleMapKind,
data_file_access_control: Arc<DataFileAccessControl>,
pub super_block: SuperBlock,
no_write_or_delete: bool,
no_write_can_delete: bool,
/// Set when a failed recovery leaves the .dat/index pair unverified: all
/// I/O is refused and a `.unavailable` marker keeps the volume quarantined
/// across restarts. Mirrors Go's ioUnavailable.
io_unavailable: Arc<IoUnavailable>,
/// Shared flag from the parent DiskLocation indicating low disk space.
/// Matches Go's `v.location.isDiskSpaceLow` checked in `IsReadOnly()`.
pub location_disk_space_low: Arc<AtomicBool>,
last_modified_ts_seconds: u64,
last_append_at_ns: u64,
pub last_disk_check_ns: Arc<std::sync::atomic::AtomicI64>, // for phantom volume detection cache
last_compact_index_offset: u64,
last_compact_revision: u16,
/// Shared with an in-flight `CompactionJob`, which runs off the store lock.
is_compacting: Arc<AtomicBool>,
/// Compaction speed limit in bytes per second (0 = unlimited).
pub compaction_byte_per_second: i64,
/// Consecutive storage-media errors and the quarantine they lead to,
/// for volume health monitoring.
io_errors: Arc<IoErrorTracker>,
/// Protobuf VolumeInfo for tiered storage (.vif file).
///
/// Private: `volume_info.files` and the write mode derived from it are the
/// same fact, so outside callers edit the list through
/// [`Volume::update_remote_files`] and read it through
/// [`Volume::volume_info`].
volume_info: PbVolumeInfo,
}
/// What a volume is created with beyond its id, directories and index kind:
/// the tail of Go's `NewVolume` argument list. The default is an empty
/// collection with no replication, no TTL and no preallocation at the
/// current needle version, which is what most tests want.
#[derive(Clone, Copy, Debug)]
pub struct VolumeSpec<'a> {
pub collection: &'a str,
pub replica_placement: Option<ReplicaPlacement>,
pub ttl: Option<crate::storage::needle::ttl::TTL>,
/// Bytes to reserve for the .dat up front; 0 grows on demand.
pub preallocate: u64,
pub version: Version,
}
impl Default for VolumeSpec<'_> {
fn default() -> Self {
VolumeSpec {
collection: "",
replica_placement: None,
ttl: None,
preallocate: 0,
version: Version::current(),
}
}
}
impl Volume {
/// Create and load a volume from disk.
pub fn new(
dirname: &str,
dir_idx: &str,
id: VolumeId,
needle_map_kind: NeedleMapKind,
spec: &VolumeSpec<'_>,
) -> Result<Self, VolumeError> {
let VolumeSpec {
collection,
replica_placement,
ttl,
preallocate,
version,
} = *spec;
let mut v = Volume {
id,
dir: dirname.to_string(),
dir_idx: dir_idx.to_string(),
collection: collection.to_string(),
dat_file: None,
remote_dat_file: None,
#[cfg(test)]
fail_fsync_for_test: false,
#[cfg(test)]
fail_idx_sync_for_test: false,
#[cfg(test)]
fail_truncate_for_test: false,
#[cfg(test)]
fail_append_for_test: HashSet::new(),
#[cfg(test)]
dat_syncs_for_test: Default::default(),
#[cfg(test)]
idx_syncs_for_test: Default::default(),
nm: None,
needle_map_kind,
data_file_access_control: Arc::new(DataFileAccessControl::default()),
super_block: SuperBlock {
replica_placement: replica_placement.unwrap_or_default(),
ttl: ttl.unwrap_or(crate::storage::needle::ttl::TTL::EMPTY),
..SuperBlock::default()
},
no_write_or_delete: false,
no_write_can_delete: false,
io_unavailable: Arc::default(),
location_disk_space_low: Arc::new(AtomicBool::new(false)),
last_modified_ts_seconds: 0,
last_append_at_ns: 0,
last_disk_check_ns: Arc::new(std::sync::atomic::AtomicI64::new(0)),
last_compact_index_offset: 0,
last_compact_revision: 0,
is_compacting: Arc::new(AtomicBool::new(false)),
compaction_byte_per_second: 0,
io_errors: Arc::default(),
volume_info: PbVolumeInfo::default(),
};
v.load(true, true, preallocate, version)?;
Ok(v)
}
/// Build a minimal, unloaded Volume holding only the identity/path fields
/// reconcile_compact_state needs. Used by the disk-location load pre-pass to
/// recover an interrupted commit before the volume itself is loaded.
pub fn new_unloaded(dirname: &str, dir_idx: &str, collection: &str, id: VolumeId) -> Self {
Volume {
id,
dir: dirname.to_string(),
dir_idx: dir_idx.to_string(),
collection: collection.to_string(),
dat_file: None,
remote_dat_file: None,
#[cfg(test)]
fail_fsync_for_test: false,
#[cfg(test)]
fail_idx_sync_for_test: false,
#[cfg(test)]
fail_truncate_for_test: false,
#[cfg(test)]
fail_append_for_test: HashSet::new(),
#[cfg(test)]
dat_syncs_for_test: Default::default(),
#[cfg(test)]
idx_syncs_for_test: Default::default(),
nm: None,
needle_map_kind: NeedleMapKind::InMemory,
data_file_access_control: Arc::new(DataFileAccessControl::default()),
super_block: SuperBlock::default(),
no_write_or_delete: false,
no_write_can_delete: false,
io_unavailable: Arc::default(),
location_disk_space_low: Arc::new(AtomicBool::new(false)),
last_modified_ts_seconds: 0,
last_append_at_ns: 0,
last_disk_check_ns: Arc::new(std::sync::atomic::AtomicI64::new(0)),
last_compact_index_offset: 0,
last_compact_revision: 0,
is_compacting: Arc::new(AtomicBool::new(false)),
compaction_byte_per_second: 0,
io_errors: Arc::default(),
volume_info: PbVolumeInfo::default(),
}
}
/// Identifies this volume instance: a re-created or remounted volume gets
/// a new one, even at the same compaction revision.
pub(crate) fn instance(&self) -> Arc<DataFileAccessControl> {
self.data_file_access_control.clone()
}
pub(crate) fn is_instance(&self, instance: &Arc<DataFileAccessControl>) -> bool {
Arc::ptr_eq(instance, &self.data_file_access_control)
}
/// Returns true if the volume is currently being compacted.
pub fn is_compacting(&self) -> bool {
self.is_compacting.load(Ordering::Acquire)
}
pub(crate) fn compacting_error(&self) -> VolumeError {
VolumeError::Io(io::Error::other(format!(
"volume {} is compacting",
self.id
)))
}
// ---- File naming (matching Go) ----
/// Base filename: dir/collection_id or dir/id
pub fn data_file_name(&self) -> String {
volume_file_name(&self.dir, &self.collection, self.id)
}
pub fn index_file_name(&self) -> String {
volume_file_name(&self.dir_idx, &self.collection, self.id)
}
pub fn file_name(&self, ext: &str) -> String {
match ext {
".idx" | ".cpx" | ".ldb" | ".cpldb" | ".rdb" => {
format!("{}{}", self.index_file_name(), ext)
}
_ => {
format!("{}{}", self.data_file_name(), ext)
}
}
}
pub fn version(&self) -> Version {
if self.volume_info.version != 0 {
Version(self.volume_info.version as u8)
} else {
self.super_block.version
}
}
// ---- Loading ----
fn load(
&mut self,
also_load_index: bool,
create_dat_if_missing: bool,
preallocate: u64,
version: Version,
) -> Result<(), VolumeError> {
let dat_path = self.file_name(".dat");
let mut already_has_super_block = false;
let has_volume_info_file = self.load_vif()?;
if self.volume_info.read_only && !self.has_remote_file() {
if self.volume_info.read_only_can_delete {
self.no_write_can_delete = true;
} else {
self.no_write_or_delete = true;
}
}
if self.has_remote_file() {
self.load_remote_dat_file()?;
if let Some(remote_file) = self.volume_info.files.first() {
if remote_file.modified_time > 0 {
self.last_modified_ts_seconds = remote_file.modified_time;
} else if let Ok(metadata) = fs::metadata(self.vif_path()) {
self.last_modified_ts_seconds = metadata
.modified()
.unwrap_or(SystemTime::UNIX_EPOCH)
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_secs();
}
}
already_has_super_block = true;
} else if Path::new(&dat_path).exists() {
let metadata = fs::metadata(&dat_path)?;
// Try to open read-write; fall back to read-only
match open_volume_file(OpenOptions::new().read(true).write(true), &dat_path) {
Ok(file) => {
self.dat_file = Some(file);
}
Err(e) if e.kind() == io::ErrorKind::PermissionDenied => {
self.dat_file =
Some(open_volume_file(OpenOptions::new().read(true), &dat_path)?);
self.no_write_or_delete = true;
}
Err(e) => return Err(e.into()),
}
self.last_modified_ts_seconds = metadata
.modified()
.unwrap_or(SystemTime::UNIX_EPOCH)
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_secs();
if metadata.len() >= SUPER_BLOCK_SIZE as u64 {
already_has_super_block = true;
}
} else if create_dat_if_missing {
// Create directory if needed
if let Some(parent) = Path::new(&dat_path).parent() {
fs::create_dir_all(parent)?;
}
let file = open_volume_file(
OpenOptions::new().read(true).write(true).create(true),
&dat_path,
)?;
if preallocate > 0 {
preallocate_file(&file, preallocate);
}
self.dat_file = Some(file);
} else {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::NotFound,
format!("volume data file {} does not exist", dat_path),
)));
}
if already_has_super_block {
match self.read_super_block() {
Ok(()) => {
if !self.super_block.version.is_supported() {
return Err(VolumeError::UnsupportedVersion(self.super_block.version.0));
}
// Match Go: v.volumeInfo.Version = uint32(v.SuperBlock.Version)
self.volume_info.version = self.super_block.version.0 as u32;
}
Err(e) if self.has_remote_file() => {
warn!(
volume_id = self.id.0,
error = %e,
"failed to read remote super block during load"
);
}
Err(e) => return Err(e),
}
} else {
self.maybe_write_super_block(version)?;
}
if also_load_index {
// Adjust for existing volumes with .idx together with .dat files:
// an index already beside the data keeps serving after --dir.idx
// named a different directory.
if self.dir_idx != self.dir
&& Path::new(&format!("{}.idx", self.data_file_name())).exists()
{
self.dir_idx = self.dir.clone();
}
// A changed --dir.idx leaves the new directory without an index.
// The .dat still holds every row, so rebuild rather than mount the
// volume with every needle invisible.
if !self.has_remote_file()
&& !Path::new(&self.file_name(".idx")).exists()
&& self.current_dat_file_size()? > SUPER_BLOCK_SIZE as u64
{
self.rebuild_idx_file()?;
info!(
volume_id = self.id.0,
idx = %self.file_name(".idx"),
"rebuilt the index from the data file"
);
}
// Recover rows that deletes on a tiered read-only volume overwrote
// at the front of .idx. Best effort: a volume that cannot be
// repaired is still servable for everything the surviving rows
// index.
match self.repair_idx_head_tombstones() {
Ok(0) => {}
Ok(restored) => info!(
volume_id = self.id.0,
restored, "recovered overwritten .idx rows from .dat"
),
Err(e) => warn!(
volume_id = self.id.0,
error = %e,
"recover overwritten .idx rows"
),
}
self.load_index()?;
// Match Go: CheckVolumeDataIntegrity after loading index (volume_loading.go L154-159)
// Only for non-remote volumes (remote storage may not have local .dat)
if !self.has_remote_file() {
if let Err(e) = self.check_volume_data_integrity() {
self.no_write_or_delete = true;
warn!(
volume_id = self.id.0,
error = %e,
"volumeDataIntegrityChecking failed"
);
}
self.recover_last_modified_ts();
// Structural check: no .idx entry may reference bytes past the
// end of .dat. The needle map's load walk above already
// populated max_needle_end, so this is a numeric comparison
// — no extra disk I/O. A violation marks the volume read-only
// so vacuum doesn't silently drop reachable data based on a
// corrupt .idx left over from a crashed batched write.
if let Some(ref nm) = self.nm
&& let Ok(dat_size) = self.current_dat_file_size()
{
let max_end = nm.max_needle_end();
if dat_size > 0 && max_end > dat_size as i64 {
self.no_write_or_delete = true;
warn!(
volume_id = self.id.0,
max_needle_end = max_end,
dat_size,
"idx references bytes past end of .dat; marking volume read-only"
);
}
}
}
}
self.restore_unavailable();
// Match Go: if no .vif file existed, create one with version and bytes_offset
if !has_volume_info_file {
self.volume_info.version = self.super_block.version.0 as u32;
self.volume_info.bytes_offset = OFFSET_SIZE as u32;
if let Err(e) = self.save_volume_info() {
warn!(
volume_id = self.id.0,
error = %e,
"failed to save volume info"
);
}
}
Ok(())
}
fn nm_or_not_found(&self) -> Result<&NeedleMap, VolumeError> {
self.nm.as_ref().ok_or_else(|| {
tracing::warn!(volume_id = self.id.0, "needle map not loaded");
VolumeError::NotFound
})
}
fn load_index(&mut self) -> Result<(), VolumeError> {
let use_redb = matches!(
self.needle_map_kind,
NeedleMapKind::Redb | NeedleMapKind::RedbMedium | NeedleMapKind::RedbLarge
);
let idx_path = self.file_name(".idx");
// Ensure idx directory exists
if let Some(parent) = Path::new(&idx_path).parent() {
fs::create_dir_all(parent)?;
}
// Go picks the sorted-file map for both read-only modes, which is every
// cloud-tiered volume: the index stays on disk in .sdx instead of in
// RAM, and nothing is held open between lookups.
if self.use_sorted_index() {
self.load_index_sorted_file(&idx_path)?;
} else if use_redb {
self.load_index_redb(&idx_path)?;
} else {
self.load_index_inmemory(&idx_path)?;
}
Ok(())
}
/// Whether this volume's index belongs in a `.sdx`. Mirrors Go's
/// `v.noWriteOrDelete || v.noWriteCanDelete` branch in volume_loading.go —
/// a tiered volume is always the second of the two.
fn use_sorted_index(&self) -> bool {
self.no_write_or_delete || self.no_write_can_delete
}
/// Load the on-disk sorted index for a read-only or cloud-tiered volume.
fn load_index_sorted_file(&mut self, idx_path: &str) -> Result<(), VolumeError> {
if !Path::new(idx_path).exists() {
// Missing .idx with existing .dat could orphan needles
let dat_path = self.file_name(".dat");
if Path::new(&dat_path).exists() {
let dat_size = fs::metadata(&dat_path).map(|m| m.len()).unwrap_or(0);
if dat_size > SUPER_BLOCK_SIZE as u64 {
warn!(
volume_id = self.id.0,
".idx file missing but .dat exists with data; needles may be orphaned"
);
}
}
// Go opens .idx with O_CREATE only where deletes are allowed; an
// empty index yields an empty .sdx and a volume that still mounts.
// A volume that takes neither writes nor deletes must not write
// here, since its index directory can sit on a read-only mount.
if !self.no_write_or_delete {
File::create(idx_path)?;
}
}
// Building .sdx writes to the index directory. Where that cannot be
// done — a read-only mount, a full disk — keep this volume's index in
// memory rather than failing the load and taking its data offline.
match SortedFileNeedleMap::open(&self.index_file_name(), self.version()) {
Ok(nm) => {
self.nm = Some(NeedleMap::SortedFile(nm));
Ok(())
}
Err(e) => {
warn!(
volume_id = self.id.0,
error = %e,
"cannot build the sorted index; keeping this volume's index in memory"
);
match self.load_index_inmemory(idx_path) {
Ok(()) => Ok(()),
// Where deletes are allowed the in-memory loader opens .idx
// read-write, which fails on the same directory that just
// refused the .sdx. Give up the deletes rather than the
// volume: without a writer no tombstone could be recorded
// anyway, so refusing them is the honest answer, and a
// remount on a writable directory restores them.
Err(retry_err) if !self.no_write_or_delete => {
warn!(
volume_id = self.id.0,
error = %retry_err,
"cannot open the index for write; serving this volume without deletes"
);
self.no_write_or_delete = true;
// Deletes are gone with the writer, so stop advertising
// them: leaving no_write_can_delete set would report
// this volume as delete-capable in metrics and mode
// checks while every delete is refused.
self.no_write_can_delete = false;
self.load_index_inmemory(idx_path)
}
Err(retry_err) => Err(retry_err),
}
}
}
}
/// Load index using in-memory CompactNeedleMap.
fn load_index_inmemory(&mut self, idx_path: &str) -> Result<(), VolumeError> {
if self.no_write_or_delete {
// Open read-only
if Path::new(&idx_path).exists() {
let mut idx_file = open_volume_file(OpenOptions::new().read(true), idx_path)?;
let nm = CompactNeedleMap::load_from_idx(&mut idx_file, self.version())?;
self.nm = Some(NeedleMap::InMemory(nm));
} else {
// Missing .idx with existing .dat could orphan needles
let dat_path = self.file_name(".dat");
if Path::new(&dat_path).exists() {
let dat_size = fs::metadata(&dat_path).map(|m| m.len()).unwrap_or(0);
if dat_size > SUPER_BLOCK_SIZE as u64 {
warn!(
volume_id = self.id.0,
".idx file missing but .dat exists with data; needles may be orphaned"
);
}
}
self.nm = Some(NeedleMap::InMemory(CompactNeedleMap::new()));
}
} else {
// Open read-write (create if missing)
let idx_file = open_volume_file(
OpenOptions::new().read(true).write(true).create(true),
idx_path,
)?;
let idx_size = trim_torn_idx_tail(&idx_file, idx_path)?;
let mut idx_reader = io::BufReader::new(&idx_file);
let mut nm = CompactNeedleMap::load_from_idx(&mut idx_reader, self.version())?;
// Re-open for positioned writes: rows are written at
// idx_file_offset so a torn tail can be trimmed (an append-mode
// handle cannot set_len on Windows).
let write_file = OpenOptions::new().write(true).open(idx_path)?;
nm.set_idx_file(Box::new(write_file), idx_size);
self.nm = Some(NeedleMap::InMemory(nm));
}
Ok(())
}
/// Load index using disk-backed RedbNeedleMap.
fn load_index_redb(&mut self, idx_path: &str) -> Result<(), VolumeError> {
// The redb database file is stored alongside the volume files
let rdb_path = self.file_name(".rdb");
// One redb database per volume: keep its page cache small, or the
// process grows by up to redb's 1 GiB default per volume (#11179).
let cache_bytes = self.needle_map_kind.redb_cache_bytes();
if self.no_write_or_delete {
// Open read-only
if Path::new(&idx_path).exists() {
let mut idx_file = open_volume_file(OpenOptions::new().read(true), idx_path)?;
let nm = RedbNeedleMap::load_from_idx(
&rdb_path,
&mut idx_file,
self.version(),
cache_bytes,
)?;
self.nm = Some(NeedleMap::Redb(nm));
} else {
// Missing .idx with existing .dat could orphan needles
let dat_path = self.file_name(".dat");
if Path::new(&dat_path).exists() {
let dat_size = fs::metadata(&dat_path).map(|m| m.len()).unwrap_or(0);
if dat_size > SUPER_BLOCK_SIZE as u64 {
warn!(
volume_id = self.id.0,
".idx file missing but .dat exists with data; needles may be orphaned"
);
}
}
self.nm = Some(NeedleMap::Redb(RedbNeedleMap::new(&rdb_path, cache_bytes)?));
}
} else {
// Open read-write (create if missing)
let idx_file = open_volume_file(
OpenOptions::new().read(true).write(true).create(true),
idx_path,
)?;
let idx_size = trim_torn_idx_tail(&idx_file, idx_path)?;
let mut idx_reader = io::BufReader::new(&idx_file);
let mut nm = RedbNeedleMap::load_from_idx(
&rdb_path,
&mut idx_reader,
self.version(),
cache_bytes,
)?;
// Re-open for positioned writes: rows are written at
// idx_file_offset so a torn tail can be trimmed (an append-mode
// handle cannot set_len on Windows).
let write_file = OpenOptions::new().write(true).open(idx_path)?;
nm.set_idx_file(Box::new(write_file), idx_size);
self.nm = Some(NeedleMap::Redb(nm));
}
Ok(())
}
pub(crate) fn load_remote_dat_file(&mut self) -> Result<(), VolumeError> {
let (storage_name, storage_key) = self.remote_storage_name_key();
let backend = crate::remote_storage::s3_tier::global_s3_tier_registry()
.read()
.unwrap()
.get(&storage_name)
.ok_or_else(|| {
VolumeError::Io(io::Error::new(
io::ErrorKind::NotFound,
format!("remote tier backend {} not found", storage_name),
))
})?;
let remote_file = self.volume_info.files.first().ok_or_else(|| {
VolumeError::Io(io::Error::new(
io::ErrorKind::NotFound,
"remote volume has no remote file entries",
))
})?;
let file_size = if remote_file.file_size > 0 {
remote_file.file_size
} else if self.volume_info.dat_file_size > 0 {
self.volume_info.dat_file_size as u64
} else {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!("remote volume {} is missing file size metadata", self.id.0),
)));
};
self.dat_file = None;
self.remote_dat_file = Some(RemoteDatFile {
backend,
key: storage_key,
file_size,
modified_time: remote_file.modified_time,
});
Ok(())
}
pub(crate) fn read_exact_at_backend(
&self,
buf: &mut [u8],
offset: u64,
) -> Result<(), VolumeError> {
if let Some(dat_file) = self.dat_file.as_ref() {
read_exact_at(dat_file, buf, offset)?;
Ok(())
} else if let Some(remote_dat_file) = self.remote_dat_file.as_ref() {
remote_dat_file.read_exact_at(buf, offset)?;
Ok(())
} else {
Err(VolumeError::Io(io::Error::other("dat file not open")))
}
}
/// Returns true when the volume has a data backend (local .dat file or
/// remote tiered storage). Mirrors Go's `v.DataBackend != nil` check.
pub fn has_data_backend(&self) -> bool {
self.dat_file.is_some() || self.remote_dat_file.is_some()
}
pub(crate) fn current_dat_file_size(&self) -> io::Result<u64> {
if let Some(ref f) = self.dat_file {
Ok(f.metadata()?.len())
} else if let Some(ref remote_dat_file) = self.remote_dat_file {
Ok(remote_dat_file.file_size)
} else {
Ok(0)
}
}
/// Read a raw byte range from the current .dat backend.
///
/// This matches Go paths that stream directly from `DataBackend`, including
/// remote-only tiered volumes whose `.dat` is no longer present locally.
pub fn read_dat_slice(&self, offset: u64, size: usize) -> Result<Vec<u8>, VolumeError> {
let _guard = self.data_file_access_control.read_lock();
if let Some(e) = self.unavailable_error() {
return Err(e);
}
let dat_size = self.current_dat_file_size()?;
if size == 0 || offset >= dat_size {
return Ok(Vec::new());
}
let read_len = std::cmp::min(size as u64, dat_size - offset) as usize;
let mut buf = vec![0u8; read_len];
self.read_exact_at_backend(&mut buf, offset)?;
Ok(buf)
}
/// Clone the remote backend handle (cheap: an `Arc` plus key/size) so a
/// caller can stream a tiered `.dat` from S3 after dropping the store lock.
/// Returns `None` for local volumes. `has_remote_file()` implies `dat_file`
/// is `None`, so this selects the same backend `read_dat_slice` would.
pub(crate) fn remote_dat_file(&self) -> Option<RemoteDatFile> {
self.remote_dat_file.clone()
}
// ---- SuperBlock I/O ----
fn read_super_block(&mut self) -> Result<(), VolumeError> {
let mut header = [0u8; SUPER_BLOCK_SIZE];
self.read_exact_at_backend(&mut header, 0)?;
let extra_size = u16::from_be_bytes([header[6], header[7]]);
let total_size = SUPER_BLOCK_SIZE + extra_size as usize;
let mut full_buf = vec![0u8; total_size];
full_buf[..SUPER_BLOCK_SIZE].copy_from_slice(&header);
if extra_size > 0 {
self.read_exact_at_backend(&mut full_buf[SUPER_BLOCK_SIZE..], SUPER_BLOCK_SIZE as u64)?;
}
self.super_block = SuperBlock::from_bytes(&full_buf)?;
// Match Go: if volumeInfo.Replication is set, override super block's ReplicaPlacement
if !self.volume_info.replication.is_empty() {
let rp = ReplicaPlacement::from_string(&self.volume_info.replication)?;
self.super_block.replica_placement = rp;
}
Ok(())
}
fn maybe_write_super_block(&mut self, version: Version) -> Result<(), VolumeError> {
let dat_file = self
.dat_file
.as_mut()
.ok_or_else(|| VolumeError::Io(io::Error::other("dat file not open")))?;
let dat_size = dat_file.metadata()?.len();
if dat_size == 0 {
if !version.is_supported() {
return Err(VolumeError::UnsupportedVersion(version.0));
}
self.super_block.version = version;
let bytes = self.super_block.to_bytes();
dat_file.seek(SeekFrom::Start(0))?;
dat_file.write_all(&bytes)?;
dat_file.sync_all()?;
}
Ok(())
}
// ---- Read ----
/// Read a needle by its ID from the volume.
pub fn read_needle(&self, n: &mut Needle) -> Result<i32, VolumeError> {
let mut read_option = ReadOption::default();
self.read_needle_with_option(n, &mut read_option)
}
pub fn read_needle_opt(&self, n: &mut Needle, read_deleted: bool) -> Result<i32, VolumeError> {
let mut read_option = ReadOption {
read_deleted,
..ReadOption::default()
};
self.read_needle_with_option(n, &mut read_option)
}
pub fn read_needle_with_option(
&self,
n: &mut Needle,
read_option: &mut ReadOption,
) -> Result<i32, VolumeError> {
let _guard = self.data_file_access_control.read_lock();
if let Some(e) = self.unavailable_error() {
return Err(e);
}
let nm = self.nm_or_not_found()?;
let nv = nm.get(n.id)?.ok_or(VolumeError::NotFound)?;
if nv.offset.is_zero() {
return Err(VolumeError::NotFound);
}
let mut read_size = nv.size;
if read_size.is_deleted() {
if read_option.read_deleted && !read_size.is_tombstone() {
// Negate to get original size
read_size = Size(-read_size.0);
} else {
return Err(VolumeError::Deleted);
}
}
if read_size.0 == 0 {
return Ok(0);
}
match self.read_needle_data_at_unlocked(
n,
nv.offset.to_actual_offset(),
read_size,
read_option,
) {
Ok(()) => self.check_read_write_error(None),
Err(VolumeError::Io(ref e)) => {
self.check_read_write_error(Some(e));
return Err(VolumeError::Io(io::Error::new(e.kind(), e.to_string())));
}
Err(e) => return Err(e),
}
if needle_expired(n) {
return Err(VolumeError::NotFound);
}
Ok(n.data_size as i32)
}
/// Read needle data from .dat file at given offset.
pub fn read_needle_data_at(
&self,
n: &mut Needle,
offset: i64,
size: Size,
) -> Result<(), VolumeError> {
let _guard = self.data_file_access_control.read_lock();
if let Some(e) = self.unavailable_error() {
return Err(e);
}
let mut read_option = ReadOption::default();
self.read_needle_data_at_unlocked(n, offset, size, &mut read_option)
}
fn read_needle_data_at_unlocked(
&self,
n: &mut Needle,
offset: i64,
size: Size,
_read_option: &mut ReadOption,
) -> Result<(), VolumeError> {
read_needle_with(
|buf, at| self.read_exact_at_backend(buf, at),
n,
offset,
size,
self.version(),
)
}
/// Read raw needle blob at a specific offset.
pub fn read_needle_blob(&self, offset: i64, size: Size) -> Result<Vec<u8>, VolumeError> {
let _guard = self.data_file_access_control.read_lock();
if let Some(e) = self.unavailable_error() {
return Err(e);
}
self.read_needle_blob_unlocked(offset, size)
}
fn read_needle_blob_unlocked(&self, offset: i64, size: Size) -> Result<Vec<u8>, VolumeError> {
let version = self.version();
// Storage guard: negativity-only (Go parity). See parse_needle_at.
if size.0 < 0 {
return Err(VolumeError::Io(std::io::Error::new(
std::io::ErrorKind::InvalidData,
format!("invalid needle size {}", size.0),
)));
}
let actual_size = get_actual_size(size, version);
let mut buf = vec![0u8; actual_size as usize];
self.read_exact_at_backend(&mut buf, offset as u64)?;
Ok(buf)
}
/// Read needle metadata at a specific offset without loading the data payload.
///
/// Matches Go's `readNeedleMetaAt`, including the tombstone path where a
/// deleted idx entry passes a negative size and the tombstone record itself
/// is read as size 0 metadata.
pub fn read_needle_meta_at(
&self,
n: &mut Needle,
offset: i64,
size: Size,
) -> Result<(), VolumeError> {
let _guard = self.data_file_access_control.read_lock();
if let Some(e) = self.unavailable_error() {
return Err(e);
}
self.read_needle_meta_at_unlocked(n, offset, size)
}
fn read_needle_meta_at_unlocked(
&self,
n: &mut Needle,
offset: i64,
size: Size,
) -> Result<(), VolumeError> {
let normalized_size = if size.is_deleted() { Size(0) } else { size };
let read_at = |buf: &mut [u8], at: u64| self.read_exact_at_backend(buf, at);
let version = self.version();
let mut offset = offset;
with_offset_retry(&mut offset, |off| {
Self::read_needle_meta_blob_and_parse(read_at, n, off, normalized_size, version)
})
}
fn read_needle_meta_blob_and_parse(
read_at: impl Fn(&mut [u8], u64) -> Result<(), VolumeError>,
n: &mut Needle,
offset: i64,
size: Size,
version: Version,
) -> Result<(), VolumeError> {
// Step 1: Read only the first 20 bytes (header + DataSize).
// Matches Go's ReadNeedleMeta which reads NeedleHeaderSize+DataSizeSize first.
const HEADER_PREFIX: usize = NEEDLE_HEADER_SIZE + DATA_SIZE_SIZE; // 20
let mut header_buf = [0u8; HEADER_PREFIX];
#[cfg(test)]
needle_read_hook::fire(n.id, header_buf.len());
read_at(&mut header_buf, offset as u64)?;
// Parse header to get the needle's Size field for validation
let (_, _, found_size) = Needle::parse_header(&header_buf);
if found_size != size {
return Err(VolumeError::Needle(NeedleError::SizeMismatch {
offset,
id: n.id,
found: found_size,
expected: size,
}));
}
// Step 2: Calculate how much meta tail to read (skip the data payload)
let actual_size = get_actual_size(size, version);
if size.0 == 0 || version == VERSION_1 {
// Tombstone or V1: no body data section, tail starts right after header
let meta_size = actual_size - NEEDLE_HEADER_SIZE as i64;
if !(0..=128 * 1024).contains(&meta_size) {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"invalid needle meta size {}: DataSize=0, size={}, offset={}",
meta_size, size.0, offset
),
)));
}
let mut meta_buf = vec![0u8; meta_size as usize];
#[cfg(test)]
needle_read_hook::fire(n.id, meta_buf.len());
read_at(&mut meta_buf, (offset + NEEDLE_HEADER_SIZE as i64) as u64)?;
n.read_paged_meta(&header_buf, &meta_buf, offset, size, version)?;
} else {
// V2/V3: extract DataSize from bytes 16..20
let data_size = u32::from_be_bytes([
header_buf[NEEDLE_HEADER_SIZE],
header_buf[NEEDLE_HEADER_SIZE + 1],
header_buf[NEEDLE_HEADER_SIZE + 2],
header_buf[NEEDLE_HEADER_SIZE + 3],
]);
// Skip past: header(16) + DataSize(4) + data(data_size)
let start_offset =
offset + NEEDLE_HEADER_SIZE as i64 + DATA_SIZE_SIZE as i64 + data_size as i64;
let stop_offset = offset + actual_size;
let meta_size = stop_offset - start_offset;
// Sanity check: reject metadata sizes > 128KB (matching Go's ReadNeedleMeta guard)
if !(0..=128 * 1024).contains(&meta_size) {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"invalid needle meta size {}: DataSize={}, size={}, offset={}",
meta_size, data_size, size.0, offset
),
)));
}
// Step 3: Read only the meta tail (skip the data payload entirely)
let mut meta_buf = vec![0u8; meta_size as usize];
#[cfg(test)]
needle_read_hook::fire(n.id, meta_buf.len());
read_at(&mut meta_buf, start_offset as u64)?;
n.read_paged_meta(&header_buf, &meta_buf, offset, size, version)?;
}
Ok(())
}
/// Resolve a needle read under the caller's store guard, without data I/O.
pub(crate) fn needle_read_plan(
&self,
id: NeedleId,
read_deleted: bool,
) -> Result<NeedleReadPlan, VolumeError> {
if let Some(e) = self.unavailable_error() {
return Err(e);
}
let nm = self.nm_or_not_found()?;
let nv = nm.get(id)?.ok_or(VolumeError::NotFound)?;
if nv.offset.is_zero() {
return Err(VolumeError::NotFound);
}
let mut size = nv.size;
if size.is_deleted() {
if read_deleted && !size.is_tombstone() {
debug!("reading deleted {}", id);
crate::metrics::HANDLER_COUNTER
.with_label_values(&[crate::metrics::READ_DELETED_NEEDLE])
.inc();
size = Size(-size.0);
} else {
return Err(VolumeError::Deleted);
}
}
if size.0 == 0 {
return Err(VolumeError::NotFound);
}
Ok(NeedleReadPlan {
source: self.read_source()?,
offset: nv.offset.to_actual_offset(),
size,
version: self.version(),
needle_id: id,
data_file_access_control: self.data_file_access_control.clone(),
io_errors: self.io_errors.clone(),
io_unavailable: self.io_unavailable.clone(),
})
}
/// A read handle on the data backend that stays valid without the guard.
fn read_source(&self) -> Result<NeedleStreamSource, VolumeError> {
if self.dat_file.is_some() {
// A fresh open, not `try_clone`: a duplicated handle shares the
// file position, and on Windows `read_exact_at` goes through
// `seek_read`, which moves it under a concurrent append. Opened
// while the caller's guard excludes a vacuum swap, the path names
// the inode `dat_file` holds.
Ok(NeedleStreamSource::Local(open_volume_file(
OpenOptions::new().read(true),
self.file_name(".dat"),
)?))
} else if let Some(remote) = self.remote_dat_file() {
Ok(NeedleStreamSource::Remote(remote))
} else {
Err(VolumeError::Io(io::Error::other("dat file not open")))
}
}
// ---- Write ----
/// Write a needle to the volume (synchronous path).
///
/// With `fsync` the write is only reported successful once both the .dat
/// record and the .idx row that indexes it are on disk, so an ack survives a
/// crash rather than sitting in the page cache. The two flushes fail
/// differently: a failed .dat flush happens before anything is published, so
/// it takes its own append back off the end and nobody else's, while a
/// failed .idx flush cannot be undone and stops the volume taking writes.
pub fn write_needle(
&mut self,
n: &mut Needle,
check_cookie: bool,
fsync: bool,
) -> Result<(u64, Size, bool), VolumeError> {
let _guard = self.data_file_access_control.write_lock();
self.check_writable()?;
self.do_write_request(n, check_cookie, fsync)
}
/// Write a batch of needles the way Go's processBatch does: one .dat
/// sync and one .idx sync per run of distinct needle ids that holds a
/// durable write, instead of two per durable needle. Nothing in such a
/// run is published before its sync, so a failed sync takes the whole
/// run back off the .dat and fails every entry. A durable entry that
/// fails to index stops the volume taking writes, and the entries after
/// it in the run are refused read only. A repeated id starts a new run,
/// so its dedup and cookie checks see the earlier write.
pub fn write_needles_grouped(
&mut self,
writes: &mut [(Needle, bool)],
) -> Vec<Result<(u64, Size, bool), VolumeError>> {
let _guard = self.data_file_access_control.write_lock();
let mut results = Vec::with_capacity(writes.len());
let mut rest = writes;
while !rest.is_empty() {
let mut ids = HashSet::new();
let len = rest.iter().take_while(|(n, _)| ids.insert(n.id)).count();
let (run, tail) = std::mem::take(&mut rest).split_at_mut(len);
if run.iter().any(|(_, fsync)| *fsync) {
results.extend(self.write_synced_run(run));
} else {
for (n, _) in run.iter_mut() {
results.push(
self.check_writable()
.and_then(|()| self.do_write_request(n, true, false)),
);
}
}
rest = tail;
}
results
}
fn write_synced_run(
&mut self,
run: &mut [(Needle, bool)],
) -> Vec<Result<(u64, Size, bool), VolumeError>> {
// Per entry: Some(offset) once appended, None when it dedups.
let mut staged = Vec::with_capacity(run.len());
// Per entry: the I/O error streak once it is staged, where a write
// sent on its own would have recorded its success.
let mut marks = Vec::with_capacity(run.len());
let mut last_append_at_ns = self.last_append_at_ns;
let mut run_start = None;
let mut sync = false;
for (n, fsync) in run.iter_mut() {
let r = self.append_unpublished(n, &mut last_append_at_ns);
marks.push(self.io_errors.mark());
if let Ok(Some(offset)) = r {
run_start.get_or_insert(offset);
}
// Only a durable entry that got this far needs the sync.
sync |= *fsync && r.is_ok();
staged.push(r);
}
if sync && let Err(e) = self.flush_dat() {
self.check_read_write_error(Some(&e));
if let Some(start) = run_start {
self.undo_unsynced_append(start);
}
let e = VolumeError::Io(e);
return staged
.into_iter()
.map(|r| r.and_then(|_| Err(run_error(&e))))
.collect();
}
self.last_append_at_ns = last_append_at_ns;
// A durable entry that fails to publish stops the volume taking
// writes, so the entries after it are refused the way a lone
// write would be.
let mut refused = false;
for ((n, fsync), r) in run.iter().zip(staged.iter_mut()) {
if refused {
*r = Err(self
.check_writable()
.err()
.unwrap_or(VolumeError::ReadOnly(self.id)));
} else if let Ok(Some(offset)) = *r
&& let Err(e) = self.publish_write(n, offset, *fsync)
{
*r = Err(e);
refused = *fsync;
}
}
if sync && let Err(e) = self.flush_idx() {
return staged
.into_iter()
.map(|r| r.and_then(|_| Err(run_error(&e))))
.collect();
}
let written = run
.iter()
.zip(&staged)
.filter(|(_, r)| matches!(r, Ok(Some(_))))
.map(|((n, _), _)| n.last_modified)
.max();
let last = staged.iter().rposition(|r| matches!(r, Ok(Some(_))));
if let (Some(last_modified), Some(last)) = (written, last) {
// Sent one at a time, the last write to land would have cleared
// the streak of the appends before it, and the ones after it
// would have failed on top of that.
self.finish_write(last_modified, sync, marks[last]);
}
run.iter()
.zip(staged)
.map(|((n, _), r)| {
let size = Size(n.data_size as i32);
r.map(|staged| match staged {
Some(offset) => (offset, size, false),
None => (0, size, true),
})
})
.collect()
}
/// The checks and the append for one entry of a synced run, leaving the
/// publish to the caller. Returns the offset, or None when it dedups.
fn append_unpublished(
&mut self,
n: &mut Needle,
last_append_at_ns: &mut u64,
) -> Result<Option<u64>, VolumeError> {
self.check_writable()?;
if self.prepare_write(n, true)? {
return Ok(None);
}
n.append_at_ns = get_append_at_ns(*last_append_at_ns);
let (offset, _, _) = self.append_needle(n)?;
*last_append_at_ns = n.append_at_ns;
Ok(Some(offset))
}
fn check_writable(&self) -> Result<(), VolumeError> {
if let Some(e) = self.unavailable_error() {
return Err(e);
}
if self.is_read_only() {
return Err(VolumeError::ReadOnly(self.id));
}
Ok(())
}
/// Flush the .dat, the first half of a durable write. The .idx is flushed
/// separately by flush_idx once the row is published; the two are split so
/// nothing is indexed before the bytes it points at are down.
fn flush_dat(&self) -> io::Result<()> {
#[cfg(test)]
self.dat_syncs_for_test.fetch_add(1, Ordering::Relaxed);
#[cfg(test)]
if self.fail_fsync_for_test {
return Err(io::Error::other("injected fsync failure"));
}
match self.dat_file.as_ref() {
Some(dat_file) => dat_file.sync_all(),
None => Err(io::Error::other("dat file not open")),
}
}
/// Flush the .idx, the second half of a durable write. load() rebuilds the
/// map from .idx, not from .dat, so the row has to be down before the write
/// is acked: lose it and the volume comes back with a .dat tail the
/// integrity check cannot account for, and loads read only.
///
/// By the time this runs the row is published and cannot be taken back -
/// undoing it would mean writing to a disk that is already failing, and
/// would leave an .idx row pointing past a truncated .dat. So a failure
/// stops the volume taking writes instead: nothing more gets appended past a
/// record whose index may not survive, and the master routes writes
/// elsewhere once the volume heartbeats read only.
fn flush_idx(&mut self) -> Result<(), VolumeError> {
#[cfg(test)]
self.idx_syncs_for_test.fetch_add(1, Ordering::Relaxed);
#[cfg(test)]
if self.fail_idx_sync_for_test {
let e = io::Error::other("injected idx sync failure");
self.no_write_or_delete = true;
return Err(VolumeError::Io(e));
}
let synced = match self.nm.as_ref() {
Some(nm) => nm.sync(),
None => Ok(()),
};
if let Err(e) = synced {
self.check_read_write_error(Some(&e));
self.no_write_or_delete = true;
tracing::error!(
"volume {}: failed to flush the index after a durable write, \
marking read only: {}",
self.id.0,
e
);
return Err(VolumeError::Io(e));
}
Ok(())
}
/// Take an unflushed append back off the .dat after its sync failed.
fn truncate_dat(&self, len: u64) -> io::Result<()> {
#[cfg(test)]
if self.fail_truncate_for_test {
return Err(io::Error::other("injected truncate failure"));
}
match self.dat_file.as_ref() {
Some(dat_file) => dat_file.set_len(len),
None => Ok(()),
}
}
fn do_write_request(
&mut self,
n: &mut Needle,
check_cookie: bool,
fsync: bool,
) -> Result<(u64, Size, bool), VolumeError> {
if self.prepare_write(n, check_cookie)? {
// Nothing to append, but the write this matched may have been
// non-durable, and the caller is asking for the content to be on
// disk. Its .idx row can be sitting in the page cache too, so both
// files get flushed exactly as they would for a fresh append.
if fsync {
self.flush_dat().map_err(|e| {
self.check_read_write_error(Some(&e));
VolumeError::Io(e)
})?;
self.flush_idx()?;
}
return Ok((0, Size(n.data_size as i32), true));
}
// Update append timestamp
n.append_at_ns = get_append_at_ns(self.last_append_at_ns);
// Append to .dat file
let (offset, _body_size, _actual_size) = self.append_needle(n)?;
// Nothing is published until the bytes are down: an index entry for an
// unflushed append would resolve past the end of the file after a crash,
// and undoing it afterwards would double-count the volume's metrics.
if fsync && let Err(e) = self.flush_dat() {
self.check_read_write_error(Some(&e));
self.undo_unsynced_append(offset);
return Err(VolumeError::Io(e));
}
self.last_append_at_ns = n.append_at_ns;
self.publish_write(n, offset, fsync)?;
if fsync {
self.flush_idx()?;
}
self.finish_write(n.last_modified, fsync, self.io_errors.mark());
// Return Size(n.DataSize) as the logical size, matching Go's doWriteRequest
Ok((offset, Size(n.data_size as i32), false))
}
/// The checks a write passes before anything is appended: TTL
/// inheritance, checksum, dedup and cookie. Returns true when the needle
/// matches the stored copy and there is nothing to append.
fn prepare_write(&self, n: &mut Needle, check_cookie: bool) -> Result<bool, VolumeError> {
// TTL inheritance from volume (matching Go's writeNeedle2)
{
use crate::storage::needle::ttl::TTL;
let needle_ttl = n.ttl.unwrap_or(TTL::EMPTY);
if needle_ttl == TTL::EMPTY && self.super_block.ttl != TTL::EMPTY {
n.set_has_ttl();
n.ttl = Some(self.super_block.ttl);
}
}
// Ensure checksum is computed before dedup check
if n.checksum == crate::storage::needle::crc::CRC(0) && !n.data.is_empty() {
n.checksum = crate::storage::needle::crc::CRC::new(&n.data);
}
// Dedup check (matches Go: n.DataSize = oldNeedle.DataSize on dedup)
if let Some(old_data_size) = self.is_file_unchanged(n) {
n.data_size = old_data_size;
return Ok(true);
}
// Cookie validation for existing needle (matches Go: check whenever nm.Get returns ok)
if let Some(nm) = &self.nm
&& let Some(nv) = nm.get(n.id)?
{
let mut existing = Needle::default();
// Read only the header to check cookie
self.read_needle_header_unlocked(&mut existing, nv.offset.to_actual_offset())?;
if n.cookie.0 == 0 && !check_cookie {
n.cookie = existing.cookie;
}
if existing.cookie != n.cookie {
return Err(VolumeError::CookieMismatch(n.cookie.0));
}
}
Ok(false)
}
/// Take an append whose sync failed back off the .dat.
fn undo_unsynced_append(&mut self, offset: u64) {
if let Err(te) = self.truncate_dat(offset) {
// The rejected record is still on the end. A later append
// would bury it mid-file, where the .dat tail check cannot
// see it, so the volume fails closed instead.
self.mark_io_unavailable(format!(
"failed to truncate back to {} after a failed fsync: {}",
offset, te
));
}
}
/// Index an appended needle. `fsync` means its record is already down.
fn publish_write(&mut self, n: &Needle, offset: u64, fsync: bool) -> Result<(), VolumeError> {
// Update needle map (uses n.size = full body size, matching Go's nm.Put)
let prior = match self.nm.as_ref() {
Some(nm) => nm.get(n.id),
None => Ok(None),
};
let should_update = match prior {
Ok(Some(nv)) => (nv.offset.to_actual_offset() as u64) < offset,
Ok(None) => true,
Err(e) => {
if fsync {
// Leaves the same durable-but-unindexed record a failed put
// below does, so it gets the same treatment.
self.no_write_or_delete = true;
tracing::error!(
"volume {}: failed to read the prior mapping for a durable write at {}, \
marking read only: {}",
self.id.0,
offset,
e
);
}
return Err(VolumeError::Io(e));
}
};
if should_update {
let indexed = match self.nm.as_mut() {
Some(nm) => nm.put(n.id, Offset::from_actual_offset(offset as i64), n.size),
None => Ok(()),
};
if let Err(e) = indexed {
if fsync {
// The record is already down but nothing indexes it, the
// same state a failed .idx flush leaves behind, so it gets
// the same treatment rather than another write on top.
self.no_write_or_delete = true;
tracing::error!(
"volume {}: failed to index a durable write at {}, \
marking read only: {}",
self.id.0,
offset,
e
);
}
return Err(VolumeError::Io(e));
}
}
Ok(())
}
/// The bookkeeping after a write is fully down. `landed_at` is the point
/// in the I/O error streak where the write landed; errors counted after
/// it, by later appends of the same run, are not cleared.
fn finish_write(&mut self, last_modified: u64, idx_synced: bool, landed_at: StreakMark) {
if self.last_modified_ts_seconds < last_modified {
self.last_modified_ts_seconds = last_modified;
}
let checkpoint_ok = self.maybe_checkpoint_index(idx_synced);
// Clear the EIO streak only after the full write (data + flush +
// index + checkpoint) succeeds, so a failed fsync or checkpoint
// does not get its EIO erased by the success reset.
if checkpoint_ok {
self.io_errors.record_success_at(landed_at);
}
}
/// Take the index checkpoint the needle map asked for, data first: the
/// checkpoint makes the index durable, and an index row that outlives
/// the bytes it points at loads read-only (see load()). A failed .dat
/// flush therefore skips the checkpoint; the map keeps asking, so it is
/// retried on the next write.
///
/// When `idx_already_synced` is true the .idx has already been fsynced by
/// `flush_idx` on the fsync=true write path, so the checkpoint skips its
/// own .idx fsync to avoid a redundant one.
fn maybe_checkpoint_index(&mut self, idx_already_synced: bool) -> bool {
let due = self.nm.as_ref().is_some_and(|nm| nm.checkpoint_due());
if !due {
return true;
}
if let Err(e) = self.flush_dat() {
self.check_read_write_error(Some(&e));
tracing::warn!(
"volume {}: skipping index checkpoint, .dat flush failed: {}",
self.id.0,
e
);
return false;
}
let checkpointed = match self.nm.as_mut() {
Some(nm) => nm.checkpoint(!idx_already_synced),
None => Ok(()),
};
if let Err(e) = checkpointed {
self.check_read_write_error(Some(&e));
tracing::warn!("volume {}: index checkpoint failed: {}", self.id.0, e);
return false;
}
true
}
fn read_needle_header_unlocked(&self, n: &mut Needle, offset: i64) -> Result<(), VolumeError> {
let mut header = [0u8; NEEDLE_HEADER_SIZE];
self.read_exact_at_backend(&mut header, offset as u64)?;
n.read_header(&header);
Ok(())
}
/// Check if the needle is unchanged from the existing one on disk.
/// Returns `Some(old_data_size)` if unchanged, `None` otherwise.
/// Matches Go's isFileUnchanged which also sets n.DataSize = oldNeedle.DataSize.
fn is_file_unchanged(&self, n: &Needle) -> Option<u32> {
// Don't dedup for volumes with TTL
if self.super_block.ttl != crate::storage::needle::ttl::TTL::EMPTY {
return None;
}
if let Some(nm) = &self.nm {
let existing = match nm.get(n.id) {
Ok(existing) => existing,
Err(e) => {
// Not a lookup miss: the index could not be read. Report
// "unknown" so the caller writes the needle again rather
// than treating an unreadable index as proof of a change.
warn!(volume_id = self.id.0, error = %e, "needle map lookup failed");
return None;
}
};
if let Some(nv) = existing
&& !nv.offset.is_zero()
&& nv.size.is_valid()
{
let mut old = Needle::default();
let mut ro = ReadOption::default();
if self
.read_needle_data_at_unlocked(
&mut old,
nv.offset.to_actual_offset(),
nv.size,
&mut ro,
)
.is_ok()
&& old.cookie == n.cookie
&& old.checksum == n.checksum
&& old.data == n.data
{
return Some(old.data_size);
}
}
}
None
}
/// Append a needle to the .dat file. Returns (offset, size, actual_size).
fn append_needle(&mut self, n: &mut Needle) -> Result<(u64, Size, i64), VolumeError> {
let version = self.version();
let bytes = n.write_bytes(version);
let actual_size = bytes.len() as i64;
let dat_file = self
.dat_file
.as_mut()
.ok_or_else(|| VolumeError::Io(io::Error::other("dat file not open")))?;
let offset = dat_file.seek(SeekFrom::End(0))?;
// Check volume size limit before writing (matching Go's Append)
if offset >= MAX_POSSIBLE_VOLUME_SIZE && !n.data.is_empty() {
return Err(VolumeError::SizeLimitExceeded {
current: offset,
limit: MAX_POSSIBLE_VOLUME_SIZE,
});
}
let written = dat_file.write_all(&bytes);
#[cfg(test)]
let written = if self.fail_append_for_test.contains(&n.id) {
Err(media_error_for_test())
} else {
written
};
if let Err(e) = written {
self.undo_unsynced_append(offset);
self.check_read_write_error(Some(&e));
return Err(VolumeError::Io(e));
}
Ok((offset, n.size, actual_size))
}
// ---- Delete ----
/// Delete a needle from the volume.
pub fn delete_needle(&mut self, n: &mut Needle) -> Result<Size, VolumeError> {
let _guard = self.data_file_access_control.write_lock();
if let Some(e) = self.unavailable_error() {
return Err(e);
}
if self.no_write_or_delete {
return Err(VolumeError::ReadOnly(self.id));
}
self.do_delete_request(n)
}
fn do_delete_request(&mut self, n: &mut Needle) -> Result<Size, VolumeError> {
let (found, size, _stored_offset) = if let Some(nm) = &self.nm {
// Propagate: acknowledging the delete as a no-op because the index
// could not be read would drop the tombstone on the floor while the
// caller believes the needle is gone.
if let Some(nv) = nm.get(n.id)? {
if !nv.size.is_deleted() {
(true, nv.size, nv.offset)
} else {
(false, Size(0), Offset::default())
}
} else {
(false, Size(0), Offset::default())
}
} else {
return Ok(Size(0));
};
if !found {
return Ok(Size(0));
}
// Write tombstone: append needle with empty data
n.data = vec![];
n.append_at_ns = get_append_at_ns(self.last_append_at_ns);
let offset = if !self.has_remote_file() {
// Normal volume: append tombstone to .dat file
let (offset, _, _) = self.append_needle(n)?;
offset
} else {
// Remote-tiered volume: skip .dat append, use offset 0
0
};
self.last_append_at_ns = n.append_at_ns;
// Update index
if let Some(nm) = &mut self.nm {
nm.delete(n.id, Offset::from_actual_offset(offset as i64))?;
}
let checkpoint_ok = self.maybe_checkpoint_index(false);
// Clear the EIO streak after a successful delete (tombstone append +
// index update + checkpoint), mirroring do_write_request.
if checkpoint_ok {
self.check_read_write_error(None);
}
Ok(size)
}
// ---- Metrics ----
pub fn content_size(&self) -> u64 {
self.nm.as_ref().map_or(0, |nm| nm.content_size())
}
pub fn deleted_size(&self) -> u64 {
self.nm.as_ref().map_or(0, |nm| nm.deleted_size())
}
pub fn file_count(&self) -> i64 {
self.nm.as_ref().map_or(0, |nm| nm.file_count())
}
pub fn deleted_count(&self) -> i64 {
self.nm.as_ref().map_or(0, |nm| nm.deleted_count())
}
pub fn max_file_key(&self) -> NeedleId {
self.nm.as_ref().map_or(NeedleId(0), |nm| nm.max_file_key())
}
pub fn is_read_only(&self) -> bool {
self.no_write_or_delete
|| self.no_write_can_delete
|| self.io_unavailable.is_set()
|| self.location_disk_space_low.load(Ordering::Relaxed)
}
/// Mirrors Go's ReadOnlyReasons: `no_write_or_delete` already covers the
/// io_unavailable quarantine.
pub fn read_only_reasons(&self) -> (bool, bool, bool, bool) {
let no_write_or_delete = self.no_write_or_delete || self.io_unavailable.is_set();
let disk_space_low = self.location_disk_space_low.load(Ordering::Relaxed);
(
no_write_or_delete || self.no_write_can_delete || disk_space_low,
no_write_or_delete,
self.no_write_can_delete,
disk_space_low,
)
}
/// The reason the volume refuses all I/O, when a failed recovery left the
/// .dat/index pair unverified. Mirrors Go's unavailableError.
pub fn unavailable_error(&self) -> Option<VolumeError> {
self.io_unavailable.error()
}
/// Fail closed after a recovery could not return the volume to a verified
/// state: refuse all I/O, leave heartbeats out, and record the state so a
/// reload stays unavailable until an operator verifies the volume.
fn mark_io_unavailable(&mut self, reason: String) {
self.no_write_or_delete = true;
self.io_unavailable.set(reason.clone());
self.mark_io_quarantined();
if let Err(e) = self.persist_unavailable(&reason) {
warn!(
volume_id = self.id.0,
error = %e,
"failed to persist unavailable marker"
);
}
if let Err(e) = self.set_read_only_persist(false, true) {
warn!(
volume_id = self.id.0,
error = %e,
"failed to persist unavailable state"
);
}
error!(
volume_id = self.id.0,
"volume entered unavailable state after failed recovery: {}", reason
);
}
fn persist_unavailable(&self, reason: &str) -> io::Result<()> {
let marker = self.file_name(".unavailable");
let mut f = OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.open(&marker)?;
f.write_all(reason.as_bytes())?;
f.sync_all()?;
drop(f);
fsync_dir(&marker)
}
/// Re-arm the state a `.unavailable` marker recorded. The marker is
/// deleted manually once the .dat/index pair is verified.
fn restore_unavailable(&mut self) {
let Ok(reason) = fs::read_to_string(self.file_name(".unavailable")) else {
return;
};
self.no_write_or_delete = true;
self.io_unavailable.set(reason.trim().to_string());
self.mark_io_quarantined();
warn!(
volume_id = self.id.0,
"volume is unavailable: {}",
reason.trim()
);
}
pub fn is_no_write_or_delete(&self) -> bool {
self.no_write_or_delete
}
pub fn set_no_write_or_delete(&mut self, value: bool) {
self.no_write_or_delete = value;
}
pub fn is_no_write_can_delete(&self) -> bool {
self.no_write_can_delete
}
pub fn last_compact_revision(&self) -> u16 {
self.last_compact_revision
}
pub fn last_modified_ts(&self) -> u64 {
self.last_modified_ts_seconds
}
pub fn is_expired(&self, volume_size: u64, volume_size_limit: u64) -> bool {
if volume_size_limit == 0 {
return false;
}
if volume_size <= SUPER_BLOCK_SIZE as u64 {
return false;
}
let ttl_minutes = self.super_block.ttl.minutes();
if ttl_minutes == 0 {
return false;
}
let lived_minutes = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_secs()
.saturating_sub(self.last_modified_ts_seconds)
/ 60;
(ttl_minutes as u64) < lived_minutes
}
/// When this volume's data becomes garbage, counted from its last write.
/// Counting from the current time instead let every .vif rewrite — a
/// read-only mark, a tier upload, an EC encode — hand an already expiring
/// volume another full TTL. Zero when the volume has no TTL. Mirrors Go's
/// ExpireAtSec.
pub fn expire_at_sec(&self) -> u64 {
let ttl_seconds = self.super_block.ttl.to_seconds();
if ttl_seconds == 0 {
return 0;
}
let mut last_write_sec = self.last_modified_ts_seconds;
if last_write_sec == 0 {
last_write_sec = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_secs();
}
last_write_sec + ttl_seconds
}
pub fn is_expired_long_enough(&self, max_delay_minutes: u32) -> bool {
let ttl_minutes = self.super_block.ttl.minutes();
if ttl_minutes == 0 {
return false;
}
let removal_delay = std::cmp::min(ttl_minutes / 10, max_delay_minutes);
((ttl_minutes + removal_delay) as u64) * 60 + self.last_modified_ts_seconds
< SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_secs()
}
/// Check volume data integrity by verifying the last index entries against the .dat file.
/// Matches Go's CheckVolumeDataIntegrity (volume_checking.go L117-141).
/// Reads the last few index entries, verifies each needle header is readable and
/// consistent. On failure, marks the volume read-only.
fn check_volume_data_integrity(&mut self) -> Result<(), VolumeError> {
let idx_path = self.file_name(".idx");
if !Path::new(&idx_path).exists() {
return Ok(());
}
let idx_size = fs::metadata(&idx_path).map(|m| m.len()).unwrap_or(0) as i64;
if idx_size == 0 {
return Ok(());
}
if idx_size % NEEDLE_MAP_ENTRY_SIZE as i64 != 0 {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!("index file's size is {} bytes, maybe corrupted", idx_size),
)));
}
let version = self.version();
// The deeper-than-tail structural check (every (offset + actual size)
// fits inside .dat) is now handled in load() via the
// needle map's max_needle_end accumulator, so we don't pay for a
// second linear scan of the .idx here.
// Verify the needle physically last in the .dat (the highest offset)
// and that the .dat ends exactly at it. The .idx is not always in .dat
// append order: weed fix and other rebuilds could rewrite it sorted by
// key, which puts the highest-key needle last instead of the .dat-tail
// needle. Picking the last file-position entry there compares a
// mid-file needle's tail against the full .dat size and falsely flips
// the volume read-only on every load. Mirrors Go's
// findDatTailEntryOffset + doCheckAndFixVolumeData.
let mut idx_file = File::open(&idx_path)?;
let tail_entry_pos = self.find_dat_tail_entry_offset(&mut idx_file, idx_size, version)?;
if tail_entry_pos < 0 {
// pre-allocated / all-zero index, nothing to verify against the .dat
return Ok(());
}
let mut buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
idx_file.seek(SeekFrom::Start(tail_entry_pos as u64))?;
idx_file.read_exact(&mut buf)?;
let (key, offset, size) = idx_entry_from_bytes(&buf);
if offset.is_zero() {
return Ok(());
}
if size.is_deleted() {
// Deletion tombstone: the .idx carries TombstoneFileSize (-1) but
// the appended needle in .dat carries Size=0, so verify with
// Size=0 — comparing against -1 would mismatch the on-disk header
// and send every volume with a trailing deletion read-only.
self.verify_deleted_needle_integrity(offset.to_actual_offset(), key, version)
} else {
self.verify_needle_integrity(offset.to_actual_offset(), key, size, version)
}
}
/// Bounds the work a vacuumed volume's recovery does, where key order makes
/// every write a candidate for the newest one. A volume with more live
/// needles than this keeps the .dat mtime: reading a subset could recover a
/// timestamp older than the newest write and expire data still inside its
/// TTL, so a scan that will not fit declines instead of guessing.
const VACUUMED_LAST_WRITE_SCAN_ENTRIES: usize = 1 << 16;
/// Point the TTL clock at the newest write recorded in the volume, replacing
/// the .dat mtime the loader starts from. A delete appends a tombstone and
/// vacuum rewrites the .dat wholesale, so the mtime moves without any write
/// ever landing: every restart of a volume taking delete traffic re-armed
/// is_expired() for another full TTL and the volume was never reclaimed.
/// Mirrors Go's recoverLastModifiedTs.
fn recover_last_modified_ts(&mut self) {
if self.super_block.ttl.minutes() == 0 {
return;
}
match self.find_last_write_append_at_ns() {
Ok(0) => {}
Ok(append_at_ns) => self.last_modified_ts_seconds = append_at_ns / 1_000_000_000,
Err(e) => warn!(
volume_id = self.id.0,
error = %e,
"recover the last write from the index"
),
}
}
/// Scan the .idx backwards for the newest write — an entry that is not a
/// deletion tombstone — and return that needle's append timestamp. The .idx
/// and the .dat share an order, so an append-ordered volume answers with the
/// first write the scan reaches. Vacuum rewrites both in key order, which
/// tracks write order only because the master issues keys increasing: an
/// overwrite keeps its original, lower key, so a vacuumed volume has to take
/// the maximum over every write it indexes. Returns 0 when the .idx holds
/// nothing but tombstones, when a vacuumed volume holds more needles than
/// the scan budget, or for a volume older than version 3, whose needles
/// carry no append timestamp. Mirrors Go's findLastWriteAppendAtNs.
fn find_last_write_append_at_ns(&self) -> Result<u64, VolumeError> {
let version = self.version();
if version != VERSION_3 {
return Ok(0);
}
let idx_path = self.file_name(".idx");
let idx_size = fs::metadata(&idx_path).map(|m| m.len()).unwrap_or(0) as i64;
if idx_size == 0 || idx_size % NEEDLE_MAP_ENTRY_SIZE as i64 != 0 {
return Ok(0);
}
let scan_every_write = self.super_block.compaction_revision > 0;
let mut entry_budget = Self::VACUUMED_LAST_WRITE_SCAN_ENTRIES;
let mut last_write_append_at_ns = 0u64;
let mut idx_file = File::open(&idx_path)?;
let mut block = vec![0u8; NEEDLE_MAP_ENTRY_SIZE * idx::ROWS_TO_READ];
let mut end = idx_size;
while end > 0 {
let start = (end - block.len() as i64).max(0);
let entries = &mut block[..(end - start) as usize];
idx_file.seek(SeekFrom::Start(start as u64))?;
idx_file.read_exact(entries)?;
for entry in entries.as_chunks::<NEEDLE_MAP_ENTRY_SIZE>().0.iter().rev() {
let (key, offset, size) = idx_entry_from_bytes(entry);
if offset.is_zero() || size.is_deleted() {
continue;
}
let Some(needle_offset) =
self.find_needle_offset(offset.to_actual_offset(), key, size)
else {
continue;
};
last_write_append_at_ns = last_write_append_at_ns
.max(self.read_needle_append_at_ns(needle_offset, size)?);
if !scan_every_write {
return Ok(last_write_append_at_ns);
}
entry_budget -= 1;
if entry_budget == 0 {
warn!(
volume_id = self.id.0,
budget = Self::VACUUMED_LAST_WRITE_SCAN_ENTRIES,
"too many needles to scan for the last write, keeping the .dat mtime"
);
return Ok(0);
}
}
end = start;
}
Ok(last_write_append_at_ns)
}
/// The .dat offset holding the needle an .idx entry describes, or None when
/// no needle there matches it. A .dat past MAX_POSSIBLE_VOLUME_SIZE wraps the
/// offsets in its .idx, so the needle can sit one volume size further in;
/// verify_needle_integrity retries the same way. Mirrors Go's
/// findNeedleOffset.
fn find_needle_offset(&self, actual_offset: i64, key: NeedleId, size: Size) -> Option<i64> {
for at in [
actual_offset,
actual_offset + MAX_POSSIBLE_VOLUME_SIZE as i64,
] {
let mut header = [0u8; NEEDLE_HEADER_SIZE];
if self.read_exact_at_backend(&mut header, at as u64).is_err() {
continue;
}
let (_, needle_id, needle_size) = Needle::parse_header(&header);
if needle_id == key && needle_size == size {
return Some(at);
}
}
None
}
/// Read the append timestamp a version 3 needle carries past its checksum.
fn read_needle_append_at_ns(&self, actual_offset: i64, size: Size) -> Result<u64, VolumeError> {
let ts_offset = actual_offset as u64
+ NEEDLE_HEADER_SIZE as u64
+ size.0 as u64
+ NEEDLE_CHECKSUM_SIZE as u64;
let mut ts = [0u8; TIMESTAMP_SIZE];
self.read_exact_at_backend(&mut ts, ts_offset)?;
Ok(u64::from_be_bytes(ts))
}
/// .idx file position of the entry whose needle is physically last in the
/// .dat (highest offset). The common case — an append-ordered .idx — is
/// resolved in O(1): the last entry's on-disk end equals the .dat size. A
/// key-sorted .idx (e.g. legacy weed fix output) falls back to a single
/// linear scan for the maximum offset. Returns -1 when no entry has a
/// non-zero offset (a pre-allocated / all-zero index).
fn find_dat_tail_entry_offset(
&self,
idx_file: &mut File,
idx_size: i64,
version: Version,
) -> Result<i64, VolumeError> {
if let Ok(dat_size) = self.dat_file_size()
&& dat_size > 0
{
let last_pos = idx_size - NEEDLE_MAP_ENTRY_SIZE as i64;
let mut buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
idx_file.seek(SeekFrom::Start(last_pos as u64))?;
idx_file.read_exact(&mut buf)?;
let (_, offset, size) = idx_entry_from_bytes(&buf);
if !offset.is_zero() && needle_disk_end(offset, size, version) == dat_size as i64 {
return Ok(last_pos);
}
}
// Slow path: scan the whole .idx for the entry at the maximum offset.
// Tombstone entries are not skipped: a trailing deletion is a real
// on-disk record and can be the needle physically last in the .dat.
let mut max_entry_pos: i64 = -1;
let mut max_actual_offset: i64 = -1;
let mut entry_pos: i64 = 0;
let mut buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
idx_file.seek(SeekFrom::Start(0))?;
let mut reader = io::BufReader::new(idx_file);
while entry_pos < idx_size {
reader.read_exact(&mut buf)?;
let (_, offset, _) = idx_entry_from_bytes(&buf);
if !offset.is_zero() {
let ao = offset.to_actual_offset();
if ao > max_actual_offset {
max_actual_offset = ao;
max_entry_pos = entry_pos;
}
}
entry_pos += NEEDLE_MAP_ENTRY_SIZE as i64;
}
Ok(max_entry_pos)
}
/// Verify the live needle at the .dat tail: its header must match the
/// .idx entry and, on v3, the .dat must end exactly at its on-disk end.
/// Extra bytes mean an unindexed trailing record (a torn append);
/// appending after one would place the next needle at an offset the
/// 8-byte .idx encoding may not represent, so the volume is quarantined
/// read-only instead. Mirrors Go's verifyNeedleIntegrity, whose tail
/// check is v3-only -- a v1/v2 volume Go serves read-write must not go
/// read-only here.
fn verify_needle_integrity(
&mut self,
actual_offset: i64,
key: NeedleId,
size: Size,
version: Version,
) -> Result<(), VolumeError> {
let mut header = [0u8; NEEDLE_HEADER_SIZE];
self.read_exact_at_backend(&mut header, actual_offset as u64)?;
let (_cookie, needle_id, needle_size) = Needle::parse_header(&header);
let mut checked_offset = actual_offset;
let mut checked_id = needle_id;
if needle_size != size {
// Go parity: on a size mismatch, retry past the 32GB offset
// wrap-around before giving up.
let alt_offset = actual_offset + MAX_POSSIBLE_VOLUME_SIZE as i64;
let mut alt_header = [0u8; NEEDLE_HEADER_SIZE];
self.read_exact_at_backend(&mut alt_header, alt_offset as u64)
.map_err(|_| size_mismatch_error(actual_offset, needle_id, needle_size, size))?;
let (_, alt_id, alt_size) = Needle::parse_header(&alt_header);
if alt_size != size {
return Err(size_mismatch_error(
actual_offset,
needle_id,
needle_size,
size,
));
}
checked_offset = alt_offset;
checked_id = alt_id;
}
if !key.is_empty() && checked_id != key {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"index key {:?} does not match needle Id {:?} at offset {}",
key, checked_id, checked_offset
),
)));
}
if version != VERSION_3 {
return Ok(());
}
let ts = self.read_needle_append_at_ns(checked_offset, size)?;
if ts > 0 {
self.last_append_at_ns = ts;
}
self.verify_dat_ends_at(checked_offset + get_actual_size(size, version))
}
/// Verify the deletion tombstone at the .dat tail. Tombstones are appended
/// with DataSize=0, so the on-disk header carries Size=0; verify with that
/// and then require the .dat to end exactly at the tombstone. Mirrors Go's
/// verifyDeletedNeedleIntegrity.
fn verify_deleted_needle_integrity(
&mut self,
actual_offset: i64,
key: NeedleId,
version: Version,
) -> Result<(), VolumeError> {
let record_size = get_actual_size(Size(0), version) as usize;
let mut record = vec![0u8; record_size];
self.read_exact_at_backend(&mut record, actual_offset as u64)?;
let mut n = Needle::default();
n.read_bytes(&record, actual_offset, Size(0), version)
.map_err(|e| {
VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!("read tombstone at offset {}: {}", actual_offset, e),
))
})?;
if !key.is_empty() && n.id != key {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"index key {:?} does not match needle Id {:?} at offset {}",
key, n.id, actual_offset
),
)));
}
if n.append_at_ns > 0 {
self.last_append_at_ns = n.append_at_ns;
}
self.verify_dat_ends_at(actual_offset + get_actual_size(Size(0), version))
}
/// The checked needle is the one physically last in the .dat, so the file
/// must end exactly at it; extra bytes mean an unindexed trailing record.
fn verify_dat_ends_at(&self, file_tail_offset: i64) -> Result<(), VolumeError> {
let file_size = self.dat_file_size().map_err(VolumeError::Io)? as i64;
if file_size == file_tail_offset {
return Ok(());
}
let dat_path = self.file_name(".dat");
if file_size > file_tail_offset {
warn!(
"data file {} actual {} bytes expected {} bytes!",
dat_path, file_size, file_tail_offset
);
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"data file {} actual {} bytes expected {} bytes",
dat_path, file_size, file_tail_offset
),
)));
}
warn!(
"data file {} has {} bytes, less than expected {} bytes!",
dat_path, file_size, file_tail_offset
);
Err(VolumeError::Io(io::Error::new(
io::ErrorKind::UnexpectedEof,
format!(
"data file {} has {} bytes, less than expected {} bytes",
dat_path, file_size, file_tail_offset
),
)))
}
/// Open the on-disk .idx for scrubbing and apply Go openIndex's zero-size
/// guard: a zero-size index is only legal for a pre-allocated volume whose
/// .dat is superblock-only; an empty index over a populated .dat is
/// corruption. Returns the open file + size, or the messages to report.
fn open_index_for_scrub(&self) -> Result<(File, i64), Vec<String>> {
let idx_path = self.file_name(".idx");
let idx_file = File::open(&idx_path)
.map_err(|e| vec![format!("open index file {}: {}", idx_path, e)])?;
let idx_file_size = idx_file
.metadata()
.map_err(|e| vec![format!("stat index file {}: {}", idx_path, e)])?
.len() as i64;
if idx_file_size == 0 {
match self.dat_file_size() {
Ok(dat_size) if dat_size > SUPER_BLOCK_SIZE as u64 => {
return Err(vec![format!(
"zero-size IDX file for volume {} with store size {}",
self.id.0, dat_size
)]);
}
Ok(_) => {}
Err(e) => {
return Err(vec![format!(
"stat data file for volume {}: {}",
self.id.0, e
)]);
}
}
}
Ok((idx_file, idx_file_size))
}
/// Scrub the volume index: verify the on-disk .idx for overlapping needles
/// and a size that is a whole number of entries. Mirrors Go's Volume.ScrubIndex
/// → idx.CheckIndexFile. Index-only; does not read the .dat.
pub fn scrub_index(&self) -> Result<(u64, Vec<String>), VolumeError> {
let _guard = self.data_file_access_control.read_lock();
let (mut idx_file, idx_file_size) = match self.open_index_for_scrub() {
Ok(v) => v,
Err(msgs) => return Ok((0, msgs)),
};
Ok(crate::storage::idx::check_index_file(
&mut idx_file,
idx_file_size,
self.version(),
))
}
/// Scrub the volume: walk the on-disk .idx (every physical row, including
/// superseded and deleted entries), CRC-verify each live needle, and reconcile
/// the .dat size against the needles it indexes. Mirrors Go's Volume.Scrub →
/// scrubVolumeData. Returns (rows_walked, broken_needles).
pub fn scrub(&self) -> Result<(u64, Vec<String>), VolumeError> {
// Matches Go's scrubVolumeData: a volume with no data backend can't be
// scrubbed (and a zero dat_size would flag every needle out of range).
if !self.has_data_backend() {
return Err(VolumeError::NotFound);
}
let _guard = self.data_file_access_control.read_lock();
let (mut idx_file, idx_file_size) = match self.open_index_for_scrub() {
Ok(v) => v,
Err(msgs) => return Ok((0, msgs)),
};
let version = self.version();
let dat_size = self.dat_file_size().map_err(VolumeError::Io)?;
// Full scrub also runs the index structural check.
let (_, mut broken) =
crate::storage::idx::check_index_file(&mut idx_file, idx_file_size, version);
let mut count: u64 = 0;
let mut total_read: i64 = 0;
let walk =
crate::storage::idx::walk_index_file(&mut idx_file, 0, |needle_id, offset, size| {
count += 1;
// A remote-tier delete records an offset-0 tombstone with no physical
// .dat bytes, so it must not contribute to total_read.
if offset.is_zero() && size.is_deleted() {
return Ok(());
}
let on_disk_size = Size(size.raw() as i32);
// compute the actual size of the needle in disk, including needle header, body and alignment padding.
total_read += get_actual_size(on_disk_size, version);
let actual_offset = offset.to_actual_offset();
if actual_offset < 0 || actual_offset as u64 >= dat_size {
broken.push(format!(
"needle {} offset {} out of range (dat_size={})",
needle_id.0, actual_offset, dat_size
));
return Ok(());
}
// Lock held above; read directly via the unlocked path (matches Go).
let mut n = Needle {
id: needle_id,
..Needle::default()
};
let mut read_option = ReadOption::default();
if let Err(e) = self.read_needle_data_at_unlocked(
&mut n,
actual_offset,
on_disk_size,
&mut read_option,
) {
broken.push(format!(
"failed to read needle {} on volume {}: {}",
needle_id.0, self.id.0, e
));
} else if n.id != needle_id {
// The data CRC does not cover the header: a damaged needle id
// still reads clean while every live-needle lookup on this
// replica keeps finding the index key here. (Go parity:
// the check covers live needles too, not just tombstones.)
broken.push(format!(
"index key {} does not match needle's Id {} on volume {}",
needle_id.0, n.id.0, self.id.0
));
}
Ok(())
});
if let Err(e) = walk {
broken.push(format!("walk index file: {}", e));
}
// Reconcile the physical .dat size against the needles read.
let want_size = total_read + SUPER_BLOCK_SIZE as i64;
if (dat_size as i64) < want_size {
broken.push(format!(
"data file for volume {} is smaller ({}) than the {} needles it contains ({})",
self.id.0, dat_size, count, want_size
));
} else if dat_size as i64 != want_size {
broken.push(format!(
"data file size for volume {} ({}) doesn't match the size for {} needles read ({})",
self.id.0, dat_size, count, want_size
));
}
Ok((count, broken))
}
/// Capture a `.dat` scan from `from_offset` that the caller can run after
/// dropping its store guard. See `DatScanPlan` for why the offset, the
/// handle and the end bound must come from the same guard.
pub(crate) fn dat_scan_plan(&self, from_offset: u64) -> Result<DatScanPlan, VolumeError> {
if let Some(e) = self.unavailable_error() {
return Err(e);
}
Ok(DatScanPlan {
source: self.read_source()?,
version: self.version(),
from: from_offset,
end: self.current_dat_file_size()?,
owner: self.data_file_access_control.clone(),
compaction_revision: self.super_block.compaction_revision,
})
}
/// Fail unless `plan` was taken from this volume at its current
/// compaction revision.
fn check_plan_source(&self, plan: &DatScanPlan) -> Result<(), VolumeError> {
if let Some(e) = self.unavailable_error() {
return Err(e);
}
if !Arc::ptr_eq(&plan.owner, &self.data_file_access_control)
|| plan.compaction_revision != self.super_block.compaction_revision
{
return Err(VolumeError::Io(io::Error::other(format!(
"volume {} was compacted or replaced during the scan",
self.id
))));
}
Ok(())
}
/// Whether the record of `key` at `offset` in `plan` is the one the
/// needle map points at now.
pub(crate) fn is_live_in_plan(
&self,
plan: &DatScanPlan,
key: NeedleId,
offset: u64,
) -> Result<bool, VolumeError> {
self.check_plan_source(plan)?;
let nv = self.nm_or_not_found()?.get(key)?;
Ok(nv.is_some_and(|nv| nv.offset.to_actual_offset() == offset as i64))
}
/// A plan for the records appended since `plan` was taken, or `None` when
/// there are none. `next` is the offset after the last record the pass
/// visited; a pass that stopped short of its end bound is not continued.
pub(crate) fn dat_scan_plan_after(
&self,
plan: &DatScanPlan,
next: Option<u64>,
) -> Result<Option<DatScanPlan>, VolumeError> {
self.check_plan_source(plan)?;
if next.unwrap_or(plan.from) != plan.end || self.current_dat_file_size()? <= plan.end {
return Ok(None);
}
self.dat_scan_plan(plan.end).map(Some)
}
/// Insert or update a needle index entry (for low-level blob writes).
pub fn put_needle_index(
&mut self,
key: NeedleId,
offset: Offset,
size: Size,
) -> Result<(), VolumeError> {
if let Some(ref mut nm) = self.nm {
nm.put(key, offset, size).map_err(VolumeError::Io)?;
}
self.maybe_checkpoint_index(false);
Ok(())
}
/// Mark this volume as read-only (no writes or deletes) and persist.
pub fn set_read_only(&mut self) -> Result<(), VolumeError> {
self.set_read_only_persist(false, true)
}
/// Mark this volume as read-only, optionally persisting to .vif.
/// With `can_delete`, deletes still land so expiring data drains the volume.
pub fn set_read_only_persist(
&mut self,
can_delete: bool,
persist: bool,
) -> Result<(), VolumeError> {
if can_delete && !self.has_remote_file() {
// deletes append tombstones to .idx; a read-only boot attached no writer
self.attach_idx_writer_if_missing()?;
}
self.no_write_or_delete = !can_delete;
if can_delete {
self.no_write_can_delete = true;
} else if !self.has_remote_file() {
// downgrading a canDelete mark; remote volumes keep their derived flag
self.no_write_can_delete = false;
}
if persist {
self.save_vif()?;
}
Ok(())
}
/// Attach an append-only .idx writer when missing (read-only boot). Runs
/// before any flag flips so a failure cannot leave puts silently skipping
/// the on-disk append.
fn attach_idx_writer_if_missing(&mut self) -> Result<(), VolumeError> {
let needs_idx_writer = self
.nm
.as_ref()
.map(|nm| !nm.has_idx_writer())
.unwrap_or(false);
if needs_idx_writer {
let idx_path = self.file_name(".idx");
// Positioned writes: an append-mode handle cannot set_len on
// Windows.
let write_file = OpenOptions::new()
.write(true)
.create(true)
.truncate(false)
.open(&idx_path)?;
let idx_size = trim_torn_idx_tail(&write_file, &idx_path)?;
if let Some(ref mut nm) = self.nm {
nm.set_idx_file(Box::new(write_file), idx_size);
}
}
Ok(())
}
/// Mark this volume as writable (allow writes and deletes).
///
/// If the volume booted with .vif ReadOnly=true, `load_index` built the
/// needle map without an .idx writer attached, so subsequent puts would
/// silently skip the on-disk append and only mutate in-memory state —
/// surviving until the next restart, then vanishing. Re-attach a writer
/// here so writes persist again.
pub fn set_writable(&mut self) -> Result<(), VolumeError> {
if let Some(e) = self.unavailable_error() {
return Err(e);
}
let was_no_write_or_delete = self.no_write_or_delete;
let was_no_write_can_delete = self.no_write_can_delete;
self.no_write_or_delete = false;
// Remote-tiered volumes must stay no_write_can_delete regardless of marks.
if !self.has_remote_file() {
self.no_write_can_delete = false;
}
// Everything below keys off the flags above, so it runs with the volume
// already marked writable. Any failure has to put them back: a volume
// advertising writable whose needle map cannot reach .idx acknowledges
// writes that are gone after a restart.
if let Err(e) = self.publish_writable() {
self.no_write_or_delete = was_no_write_or_delete;
self.no_write_can_delete = was_no_write_can_delete;
return Err(e);
}
self.reset_io_error_state();
Ok(())
}
/// The steps that must all succeed before a volume can be left marked
/// writable: swap out the read-only map, attach the .idx writer, persist the
/// mark. Split out so set_writable has one rollback point rather than three.
fn publish_writable(&mut self) -> Result<(), VolumeError> {
self.reconcile_index_mode()?;
self.attach_idx_writer_if_missing()?;
self.save_vif()
}
/// Swap the read-only sorted map for the configured writable one once the
/// current mode says the volume is no longer read-only. Mirrors Go's
/// reopenIdxForWrite, and is the reason a volume that stops being read-only
/// can accept writes at all: SortedFileNeedleMap::put always fails, so
/// publishing a writable volume without this leaves every write appending to
/// .dat and then failing to index it. A no-op unless the map is the sorted
/// one and the flags have already moved.
fn reconcile_index_mode(&mut self) -> Result<(), VolumeError> {
let wants_sorted = self.use_sorted_index();
let has_sorted = matches!(self.nm, Some(NeedleMap::SortedFile(_)));
if wants_sorted == has_sorted {
return Ok(());
}
// Both directions matter. Tiering DOWN has to trade the read-only
// sorted map for a writable one before the volume is published as
// writable. Tiering UP has to install the sorted map, or a volume
// tiered while the server runs keeps its whole index in RAM and its
// .idx descriptor pinned until the next restart — the cost this map
// exists to avoid.
//
// load_index only publishes self.nm once the new map is built, so a
// failure leaves the previous map in place and the caller can recover.
self.load_index()
}
/// The tiered-storage VolumeInfo backing the .vif file.
pub fn volume_info(&self) -> &PbVolumeInfo {
&self.volume_info
}
/// Whether this volume is backed by a remote file, i.e. whether the .vif
/// holds a remote reference. Derived rather than mirrored, so it can never
/// disagree with the reference list it describes.
pub fn has_remote_file(&self) -> bool {
!self.volume_info.files.is_empty()
}
/// Edit the remote reference list and recompute everything derived from it.
///
/// This is the only way to change `volume_info.files` from outside the
/// module: the write/delete mode and the needle map follow from whether the
/// volume is remote, so an edit that skipped the refresh would leave the
/// volume claiming a mode its .vif contradicts.
///
/// Not transactional: on error `f` has already been applied and the volume
/// is left pinned read-only, which is the safe end of a half-finished tier
/// transition. A caller that needs the old list back restores it with a
/// second call.
pub fn update_remote_files(
&mut self,
f: impl FnOnce(&mut Vec<PbRemoteFile>),
) -> Result<(), VolumeError> {
f(&mut self.volume_info.files);
self.refresh_remote_write_mode()
}
/// Recompute the Go-style write/delete mode from the current remote tier
/// state, and bring the needle map in line with it — a volume that stops
/// being remote also stops using the read-only sorted index.
///
/// If the map cannot be rebuilt the volume is pinned read-only rather than
/// published as writable with an index that rejects every put; a restart
/// recovers it.
fn refresh_remote_write_mode(&mut self) -> Result<(), VolumeError> {
if self.has_remote_file() {
self.no_write_can_delete = true;
self.no_write_or_delete = false;
} else if !self.volume_info.read_only_can_delete {
// only clear the remoteness-derived flag, not an operator mark
self.no_write_can_delete = false;
}
if let Err(e) = self.reconcile_index_mode() {
self.no_write_or_delete = true;
return Err(e);
}
Ok(())
}
/// Open the local .dat as the data backend, dropping any remote backend, so reads
/// are served from local disk. Mirrors Go's swapToLocalDatBackend after a tier-down.
pub(crate) fn open_local_dat_backend(&mut self) -> Result<(), VolumeError> {
let file = open_volume_file(OpenOptions::new().read(true).write(true), self.dat_path())?;
self.remote_dat_file = None;
self.dat_file = Some(file);
Ok(())
}
/// Path to .vif file.
fn vif_path(&self) -> String {
format!("{}.vif", self.data_file_name())
}
/// Load volume info from .vif file.
/// Supports both the protobuf-JSON format (Go-compatible) and legacy JSON.
/// Returns true if a .vif file was found and successfully loaded.
fn load_vif(&mut self) -> Result<bool, VolumeError> {
let path = self.vif_path();
if let Ok(content) = fs::read_to_string(&path) {
if content.trim().is_empty() {
return Ok(false);
}
// Try protobuf-JSON (Go-compatible VolumeInfo via VifVolumeInfo)
if let Ok(vif_info) = serde_json::from_str::<VifVolumeInfo>(&content) {
let pb_info = vif_info.to_pb();
if pb_info.read_only {
if pb_info.read_only_can_delete {
self.no_write_can_delete = true;
} else {
self.no_write_or_delete = true;
}
}
self.volume_info = pb_info;
self.refresh_remote_write_mode()?;
if self.volume_info.version == 0 {
self.volume_info.version = Version::current().0 as u32;
}
if !self.has_remote_file() && self.volume_info.bytes_offset == 0 {
self.volume_info.bytes_offset = OFFSET_SIZE as u32;
}
if self.volume_info.bytes_offset != 0
&& self.volume_info.bytes_offset != OFFSET_SIZE as u32
{
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"bytes_offset mismatch in {}: found {}, expected {}",
path, self.volume_info.bytes_offset, OFFSET_SIZE
),
)));
}
return Ok(true);
}
// Fall back to OldVersionVolumeInfo (Go's tryOldVersionVolumeInfo):
// maps DestroyTime -> expire_at_sec
if let Ok(old_info) = serde_json::from_str::<OldVersionVifVolumeInfo>(&content) {
let vif_info = old_info.into_vif();
let pb_info = vif_info.to_pb();
if pb_info.read_only {
self.no_write_or_delete = true;
}
self.volume_info = pb_info;
self.refresh_remote_write_mode()?;
if self.volume_info.version == 0 {
self.volume_info.version = Version::current().0 as u32;
}
if !self.has_remote_file() && self.volume_info.bytes_offset == 0 {
self.volume_info.bytes_offset = OFFSET_SIZE as u32;
}
if self.volume_info.bytes_offset != 0
&& self.volume_info.bytes_offset != OFFSET_SIZE as u32
{
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"bytes_offset mismatch in {}: found {}, expected {}",
path, self.volume_info.bytes_offset, OFFSET_SIZE
),
)));
}
return Ok(true);
}
// Fall back to legacy format
if let Ok(info) = serde_json::from_str::<VolumeInfo>(&content) {
if info.read_only {
self.no_write_or_delete = true;
}
return Ok(true);
}
}
Ok(false)
}
/// Save volume info to .vif file in protobuf-JSON format (Go-compatible).
/// Matches Go's SaveVolumeInfo: checks writability before writing and propagates errors.
fn save_vif(&mut self) -> Result<(), VolumeError> {
let vif_path = self.vif_path();
// Match Go: if file exists but is not writable, return an error
let path = std::path::Path::new(&vif_path);
if path.exists() {
let metadata = fs::metadata(path)?;
if metadata.permissions().readonly() {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::PermissionDenied,
format!("failed to check {} not writable", vif_path),
)));
}
}
// remoteness-derived no_write_can_delete is not an operator mark; sync
// volume_info so refresh_remote_write_mode sees a real mark
let marked_can_delete =
self.no_write_can_delete && !self.has_remote_file() && !self.no_write_or_delete;
self.volume_info.read_only = self.no_write_or_delete || marked_can_delete;
self.volume_info.read_only_can_delete = marked_can_delete;
let mut vif = VifVolumeInfo::from_pb(&self.volume_info);
// Match Go's SaveVolumeInfo: compute ExpireAtSec from TTL
let expire_at_sec = self.expire_at_sec();
if expire_at_sec > 0 {
vif.expire_at_sec = expire_at_sec;
}
let content = serde_json::to_string_pretty(&vif)
.map_err(|e| VolumeError::Io(io::Error::other(e.to_string())))?;
fs::write(&vif_path, content)?;
Ok(())
}
/// Save full VolumeInfo to .vif file (for tiered storage).
/// Matches Go's SaveVolumeInfo which computes ExpireAtSec from TTL.
pub fn save_volume_info(&mut self) -> Result<(), VolumeError> {
let marked_can_delete =
self.no_write_can_delete && !self.has_remote_file() && !self.no_write_or_delete;
self.volume_info.read_only = self.no_write_or_delete || marked_can_delete;
self.volume_info.read_only_can_delete = marked_can_delete;
// Compute ExpireAtSec from TTL (matches Go's SaveVolumeInfo)
let expire_at_sec = self.expire_at_sec();
if expire_at_sec > 0 {
self.volume_info.expire_at_sec = expire_at_sec;
}
let vif = VifVolumeInfo::from_pb(&self.volume_info);
let content = serde_json::to_string_pretty(&vif)
.map_err(|e| VolumeError::Io(io::Error::other(e.to_string())))?;
// fsync the .vif so a tiered volume's remote reference is durable before the
// caller acts on it, e.g. deletes the remote object (matches Go util.WriteFile).
let mut f = OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.open(self.vif_path())?;
f.write_all(content.as_bytes())?;
f.sync_all()?;
Ok(())
}
/// Get the remote storage backend name and key from this volume .vif.
pub fn remote_storage_name_key(&self) -> (String, String) {
if self.volume_info.files.is_empty() {
return (String::new(), String::new());
}
let rf = &self.volume_info.files[0];
let backend_name = if rf.backend_id.is_empty() {
rf.backend_type.clone()
} else {
format!("{}.{}", rf.backend_type, rf.backend_id)
};
(backend_name, rf.key.clone())
}
/// Get the dat file path for this volume.
pub fn dat_path(&self) -> String {
self.file_name(".dat")
}
/// Get the directory this volume is stored in.
pub fn dir(&self) -> &str {
&self.dir
}
/// Get the directory this volume's index is stored in.
pub fn dir_idx(&self) -> &str {
&self.dir_idx
}
/// Throttle IO during compaction to avoid saturating disk.
pub fn maybe_throttle_compaction(&self, bytes_written: u64) {
if self.compaction_byte_per_second <= 0 || !self.is_compacting() {
return;
}
// Simple throttle: sleep based on bytes written vs allowed rate
let sleep_us =
(bytes_written as f64 / self.compaction_byte_per_second as f64 * 1_000_000.0) as u64;
if sleep_us > 0 {
std::thread::sleep(std::time::Duration::from_micros(sleep_us));
}
}
/// Change the replication placement and rewrite the super block.
pub fn set_replica_placement(&mut self, rp: ReplicaPlacement) -> Result<(), VolumeError> {
self.super_block.replica_placement = rp;
let bytes = self.super_block.to_bytes();
let dat_file = self
.dat_file
.as_mut()
.ok_or_else(|| VolumeError::Io(io::Error::other("dat file not open")))?;
dat_file.seek(SeekFrom::Start(0))?;
dat_file.write_all(&bytes)?;
dat_file.sync_all()?;
Ok(())
}
// ---- Binary search for incremental copy ----
/// Read a single index entry's offset from the .idx file by entry index.
fn read_offset_from_index(&self, m: i64) -> Result<Offset, VolumeError> {
let idx_path = self.file_name(".idx");
let idx_file = File::open(&idx_path)?;
let mut buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
let file_offset = m as u64 * NEEDLE_MAP_ENTRY_SIZE as u64;
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
idx_file.read_exact_at(&mut buf, file_offset)?;
}
#[cfg(not(unix))]
{
let mut f = idx_file;
f.seek(SeekFrom::Start(file_offset))?;
std::io::Read::read_exact(&mut f, &mut buf)?;
}
let (_key, offset, _size) = idx_entry_from_bytes(&buf);
Ok(offset)
}
/// Read the append_at_ns timestamp from a needle at the given offset in the .dat file.
/// Go reads the full needle body for ALL entries including tombstones to get the
/// actual AppendAtNs timestamp, which is needed for correct binary search during
/// incremental copy.
fn read_append_at_ns(&self, offset: Offset) -> Result<u64, VolumeError> {
let actual_offset = offset.to_actual_offset() as u64;
let version = self.version();
let mut header_buf = [0u8; NEEDLE_HEADER_SIZE];
self.read_exact_at_backend(&mut header_buf, actual_offset)?;
let (_cookie, _id, size) = Needle::parse_header(&header_buf);
let actual_size = get_actual_size(size, version);
if actual_size <= 0 {
return Ok(0);
}
let mut buf = vec![0u8; actual_size as usize];
self.read_exact_at_backend(&mut buf, actual_offset)?;
let mut n = Needle::default();
n.read_bytes_meta_only(&buf, offset.to_actual_offset(), size, version)?;
Ok(n.append_at_ns)
}
/// Search right from position m to find the first non-deleted entry.
fn read_right_ns(&self, m: i64, max: i64) -> Result<(i64, Offset, u64), VolumeError> {
let mut index = m;
loop {
index += 1;
if index >= max {
return Ok((index, Offset::default(), 0));
}
let offset = self.read_offset_from_index(index)?;
if !offset.is_zero() {
let ts = self.read_append_at_ns(offset)?;
return Ok((index, offset, ts));
}
}
}
/// Search left from position m to find the first non-deleted entry.
fn read_left_ns(&self, m: i64) -> Result<(i64, Offset, u64), VolumeError> {
let mut index = m;
loop {
index -= 1;
if index < 0 {
return Ok((index, Offset::default(), 0));
}
let offset = self.read_offset_from_index(index)?;
if !offset.is_zero() {
let ts = self.read_append_at_ns(offset)?;
return Ok((index, offset, ts));
}
}
}
/// Binary search through the .idx file to find the first needle
/// with append_at_ns > since_ns. Returns (offset, is_last).
/// Matches Go's BinarySearchByAppendAtNs in volume_backup.go.
pub fn binary_search_by_append_at_ns(
&self,
since_ns: u64,
) -> Result<(Offset, bool), VolumeError> {
let file_size = self.idx_file_size() as i64;
if file_size % NEEDLE_MAP_ENTRY_SIZE as i64 != 0 {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!("unexpected idx file size: {}", file_size),
)));
}
let entry_count = file_size / NEEDLE_MAP_ENTRY_SIZE as i64;
let mut l: i64 = 0;
let mut h: i64 = entry_count;
while l < h {
let m = (l + h) / 2;
if m == entry_count {
return Ok((Offset::default(), true));
}
let offset = self.read_offset_from_index(m)?;
if offset.is_zero() {
let (left_index, _left_offset, left_ns) = self.read_left_ns(m)?;
let (right_index, right_offset, right_ns) = self.read_right_ns(m, entry_count)?;
if right_ns <= since_ns {
l = right_index;
if l == entry_count {
return Ok((Offset::default(), true));
} else {
continue;
}
}
if since_ns < left_ns {
h = left_index + 1;
continue;
}
return Ok((right_offset, false));
}
let m_ns = self.read_append_at_ns(offset)?;
if m_ns <= since_ns {
l = m + 1;
} else {
h = m;
}
}
if l == entry_count {
return Ok((Offset::default(), true));
}
let offset = self.read_offset_from_index(l)?;
Ok((offset, false))
}
/// Write a raw needle blob at a specific offset in the .dat file.
pub fn write_needle_blob(
&mut self,
offset: i64,
needle_blob: &[u8],
) -> Result<(), VolumeError> {
if self.is_read_only() {
return Err(VolumeError::ReadOnly(self.id));
}
let dat_file = self
.dat_file
.as_mut()
.ok_or_else(|| VolumeError::Io(io::Error::other("dat file not open")))?;
dat_file.seek(SeekFrom::Start(offset as u64))?;
dat_file.write_all(needle_blob)?;
Ok(())
}
/// Write a needle blob and update the needle map index.
/// Matches Go's Volume.WriteNeedleBlob which appends to dat and calls nm.Put.
pub fn write_needle_blob_and_index(
&mut self,
needle_id: NeedleId,
needle_blob: &[u8],
size: Size,
) -> Result<(), VolumeError> {
// nm.put on a read-only volume fails only after the blob is appended to .dat.
if self.is_read_only() {
return Err(VolumeError::ReadOnly(self.id));
}
// Storage guard: negativity-only (Go parity). See parse_needle_at.
if size.0 < 0 {
return Err(VolumeError::Io(std::io::Error::new(
std::io::ErrorKind::InvalidData,
format!("invalid needle size {}", size.0),
)));
}
// size indexes the needle and places the v3 append timestamp, so a caller using
// the payload-only DataSize corrupts both, silently until the needle is read back.
if needle_blob.len() < NEEDLE_HEADER_SIZE {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"needle {} blob of {} bytes is shorter than a needle header",
needle_id.0,
needle_blob.len()
),
)));
}
let (_, _, header_size) = Needle::parse_header(needle_blob);
if header_size != size {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"needle {} size {} does not match its blob header size {}",
needle_id.0, size.0, header_size.0
),
)));
}
// The blob is appended as is: its length must be the record this size takes in this volume's version.
let actual_size = get_actual_size(size, self.version());
if needle_blob.len() as i64 != actual_size {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"needle {} blob of {} bytes does not match the {} bytes size {} takes in a version {} volume",
needle_id.0,
needle_blob.len(),
actual_size,
size.0,
self.version().0
),
)));
}
// Dedup check: if the same needle already exists with matching content, skip the write.
// Matches Go's WriteNeedleBlob which reads existing needle and compares cookie+checksum+data.
if let Some(nm) = &self.nm
&& let Some(nv) = nm.get(needle_id)?
&& nv.size == size
{
let version = self.version();
// Read existing needle from disk
let mut old_needle = Needle::default();
let mut ro = ReadOption::default();
if self
.read_needle_data_at_unlocked(
&mut old_needle,
nv.offset.to_actual_offset(),
nv.size,
&mut ro,
)
.is_ok()
{
// Parse the incoming blob into a needle
let mut new_needle = Needle::default();
if new_needle
.read_bytes(needle_blob, nv.offset.to_actual_offset(), size, version)
.is_ok()
&& old_needle.cookie == new_needle.cookie
&& old_needle.checksum == new_needle.checksum
&& old_needle.data == new_needle.data
{
return Ok(());
}
}
}
// Check volume size limit
let content_size = self.content_size();
if MAX_POSSIBLE_VOLUME_SIZE < content_size + needle_blob.len() as u64 {
return Err(VolumeError::Io(io::Error::other(format!(
"volume size limit {} exceeded! current size is {}",
MAX_POSSIBLE_VOLUME_SIZE, content_size
))));
}
// Compute monotonic appendAtNs (matches Go: needle.GetAppendAtNs(v.lastAppendAtNs))
let append_at_ns = get_append_at_ns(self.last_append_at_ns);
// Patch appendAtNs timestamp into V3 blobs (matches Go WriteNeedleBlob L64-77)
let mut blob_buf;
let blob_to_write = if self.version() == VERSION_3 {
let ts_offset = NEEDLE_HEADER_SIZE + size.0 as usize + NEEDLE_CHECKSUM_SIZE;
if ts_offset + TIMESTAMP_SIZE > needle_blob.len() {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"needle blob buffer too small: need {} bytes, have {}",
ts_offset + TIMESTAMP_SIZE,
needle_blob.len()
),
)));
}
blob_buf = needle_blob.to_vec();
blob_buf[ts_offset..ts_offset + TIMESTAMP_SIZE]
.copy_from_slice(&append_at_ns.to_be_bytes());
&blob_buf[..]
} else {
needle_blob
};
// Append blob at end of dat file
let dat_size = self.dat_file_size()? as i64;
self.write_needle_blob(dat_size, blob_to_write)?;
// Update lastAppendAtNs (matches Go L352: v.lastAppendAtNs = appendAtNs)
self.last_append_at_ns = append_at_ns;
// Update needle map index
let offset = Offset::from_actual_offset(dat_size);
if let Some(ref mut nm) = self.nm {
nm.put(needle_id, offset, size)?;
}
self.maybe_checkpoint_index(false);
Ok(())
}
pub fn needs_replication(&self) -> bool {
self.super_block.replica_placement.get_copy_count() > 1
}
/// Garbage ratio: deleted_size / content_size (matching Go's garbageLevel).
/// content_size is the additive-only FileByteCounter.
///
/// When DeletedCount > 0 but DeletedSize == 0 (e.g. .sdx converted back to
/// normal .idx where deleted entry sizes are missing), falls back to
/// computing deleted bytes as (datFileSize - contentSize - SuperBlockSize)
/// and uses datFileSize as the denominator.
pub fn garbage_level(&self) -> f64 {
let content = self.content_size();
if content == 0 {
return 0.0;
}
let mut deleted = self.deleted_size();
let mut file_size = content;
if self.deleted_count() > 0 && deleted == 0 {
// This happens for .sdx converted back to normal .idx
// where deleted entry size is missing
let dat_file_size = self.dat_file_size().unwrap_or(0);
deleted = dat_file_size
.saturating_sub(content)
.saturating_sub(SUPER_BLOCK_SIZE as u64);
file_size = dat_file_size;
}
if file_size == 0 {
return 0.0;
}
deleted as f64 / file_size as f64
}
pub fn dat_file_size(&self) -> io::Result<u64> {
self.current_dat_file_size()
}
/// Get the modification time of the .dat file as Unix seconds.
pub fn dat_file_mod_time(&self) -> u64 {
if let Some(dat_file) = self.dat_file.as_ref() {
dat_file
.metadata()
.ok()
.and_then(|m| m.modified().ok())
.and_then(|t| t.duration_since(std::time::UNIX_EPOCH).ok())
.map(|d| d.as_secs())
.unwrap_or(0)
} else {
self.remote_dat_file
.as_ref()
.map(|remote_dat_file| remote_dat_file.modified_time)
.unwrap_or(0)
}
}
pub fn idx_file_size(&self) -> u64 {
self.nm.as_ref().map_or(0, |nm| nm.index_file_size())
}
// ---- Compaction / Vacuum ----
/// Compact the volume by copying only live needles to new .cpd/.cpx files.
/// This reads from the current .dat/.idx and writes to .cpd/.cpx.
/// Call `commit_compact()` after to swap the files.
pub fn compact_by_index<F>(
&mut self,
_preallocate: u64,
_max_bytes_per_second: i64,
progress_fn: F,
) -> Result<(), VolumeError>
where
F: Fn(i64) -> bool,
{
match self.begin_compact_by_index()? {
Some(job) => job.run(progress_fn),
None => Ok(()), // already compacting
}
}
/// Claim the compaction, record the `makeup_diff` watermark, and open the
/// sources the copy reads, so the caller can run the copy after dropping
/// its store guard. `None` means a compaction is already running.
pub(crate) fn begin_compact_by_index(&mut self) -> Result<Option<CompactionJob>, VolumeError> {
let Some(claim) = CompactionClaim::try_claim(&self.is_compacting) else {
return Ok(None);
};
// Guard against nil needle map (matches Go's nil check)
let Some(nm) = self.nm.as_ref() else {
return Err(VolumeError::Io(io::Error::other(format!(
"volume {} needle map is nil",
self.id
))));
};
if let Some(e) = self.unavailable_error() {
return Err(e);
}
let idx_size = nm.index_file_size();
// The copy would stream the .dat from remote storage for a commit that refuses it.
if self.has_remote_file() {
return Err(VolumeError::Io(io::Error::other(format!(
"volume {} is tiered to remote storage, cannot compact",
self.id
))));
}
// Fresh opens, not `try_clone`: see `dat_scan_plan`.
let src_dat = if self.dat_file.is_some() {
NeedleStreamSource::Local(open_volume_file(
OpenOptions::new().read(true),
self.file_name(".dat"),
)?)
} else if let Some(remote) = self.remote_dat_file() {
NeedleStreamSource::Remote(remote)
} else {
return Err(VolumeError::Io(io::Error::other("dat file not open")));
};
let src_idx = open_volume_file(OpenOptions::new().read(true), self.file_name(".idx"))?;
// Record state before compaction for makeupDiff
self.last_compact_index_offset = idx_size;
self.last_compact_revision = self.super_block.compaction_revision;
Ok(Some(CompactionJob {
id: self.id,
src_dat,
src_idx,
idx_size,
cpd_path: self.file_name(".cpd"),
cpx_path: self.file_name(".cpx"),
version: self.version(),
super_block: self.super_block.clone(),
io_errors: self.io_errors.clone(),
_claim: claim,
}))
}
/// Commit a previously completed compaction: swap .cpd/.cpx to .dat/.idx and reload.
/// Matches Go's isCompactionInProgress CompareAndSwap guard.
pub fn commit_compact(&mut self) -> Result<(), VolumeError> {
let Some(_claim) = CompactionClaim::try_claim(&self.is_compacting) else {
return Ok(()); // already compacting, silently skip (matches Go)
};
// The reload would read the remote object through the compacted .idx.
if self.has_remote_file() {
let _ = fs::remove_file(self.file_name(".cpd"));
let _ = fs::remove_file(self.file_name(".cpx"));
return Err(VolumeError::Io(io::Error::other(format!(
"volume {} is tiered to remote storage, cannot commit compaction",
self.id
))));
}
self.do_commit_compact()
}
fn do_commit_compact(&mut self) -> Result<(), VolumeError> {
if let Err(e) = self.makeup_diff() {
warn!("makeup_diff failed: {}", e);
// Match Go: clean up .cpd/.cpx on makeup_diff failure
let cpd = self.file_name(".cpd");
let cpx = self.file_name(".cpx");
let _ = fs::remove_file(&cpd);
let _ = fs::remove_file(&cpx);
return Err(e);
}
// Close current files
if let Some(ref mut nm) = self.nm {
nm.close();
}
self.nm = None;
if let Some(ref dat_file) = self.dat_file {
let _ = dat_file.sync_all();
}
self.dat_file = None;
self.remote_dat_file = None;
// makeup_diff has fsynced the .cpd/.cpx contents. Persist a durable .cpc
// commit marker BEFORE the renames so the two-rename swap is atomic
// across a crash: a marker on disk means the swap is decided and
// reconcile rolls forward; no marker means roll back. Without it, a
// crash between the two renames leaves a stale .idx that a later vacuum
// compacts to empty.
self.write_compact_commit_marker()?;
self.apply_compact_swap()?;
// Reload
self.load(true, false, 0, self.version())?;
Ok(())
}
/// Write and fsync the .cpc marker, then fsync the directory so the marker
/// survives a crash before apply_compact_swap.
fn write_compact_commit_marker(&self) -> Result<(), VolumeError> {
let marker_path = self.file_name(".cpc");
let f = OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.open(&marker_path)?;
f.sync_all()?;
drop(f);
fsync_dir(&marker_path)?;
Ok(())
}
/// Perform the durable two-rename compaction commit. Idempotent: run both at
/// the tail of commit and by reconcile rolling forward after a crash. It
/// requires BOTH .cpd/.cpx to be present, so a stale or duplicate commit
/// returns an error without deleting the live .dat/.idx.
fn apply_compact_swap(&self) -> Result<(), VolumeError> {
// Normal commit: both temp files must be present, or a stale/duplicate
// commit could clobber the live .dat/.idx.
if !Path::new(&self.file_name(".cpd")).exists()
|| !Path::new(&self.file_name(".cpx")).exists()
{
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::NotFound,
"compact files (.cpd/.cpx) not found",
)));
}
self.finish_compact_swap()
}
/// Renames whichever compaction temp file is still present (.cpd->.dat,
/// .cpx->.idx), fsyncs, drops the stale .ldb/.rdb, then clears the .cpc
/// marker. Tolerates a partial state: a crash after the .dat rename but
/// before the .idx rename leaves only .cpx, which must still be applied --
/// not abandoned, which would pair a fresh .dat with a stale .idx. Existence
/// is checked robustly so a transient error never silently skips the swap.
fn finish_compact_swap(&self) -> Result<(), VolumeError> {
let exists = |p: String| match fs::metadata(&p) {
Ok(_) => true,
Err(e) if e.kind() == io::ErrorKind::NotFound => false,
Err(_) => true,
};
let dat_path = self.file_name(".dat");
let idx_path = self.file_name(".idx");
let cpd_exists = exists(self.file_name(".cpd"));
let cpx_exists = exists(self.file_name(".cpx"));
if cpd_exists {
#[cfg(windows)]
{
let _ = fs::remove_file(&dat_path);
}
fs::rename(self.file_name(".cpd"), &dat_path)?;
}
if cpx_exists {
#[cfg(windows)]
{
let _ = fs::remove_file(&idx_path);
}
fs::rename(self.file_name(".cpx"), &idx_path)?;
}
if cpd_exists || cpx_exists {
fsync_dir(&dat_path)?;
if self.dir != self.dir_idx {
fsync_dir(&idx_path)?;
}
let _ = fs::remove_dir_all(self.file_name(".ldb"));
let _ = fs::remove_file(self.file_name(".rdb"));
}
// Clear the marker last and fsync the dir so a restart does not re-run a
// completed swap.
let marker_path = self.file_name(".cpc");
match fs::remove_file(&marker_path) {
Ok(()) => {}
Err(e) if e.kind() == io::ErrorKind::NotFound => {}
Err(e) => return Err(e.into()),
}
fsync_dir(&marker_path)?;
Ok(())
}
/// Recover an interrupted compaction commit on load. When the .cpc marker is
/// present the swap was decided, so roll FORWARD by finishing the renames;
/// when it is absent any leftover .cpd/.cpx are an abandoned generation, so
/// roll BACK by deleting them (and any stale .ldb beside a healthy .idx).
pub fn reconcile_compact_state(&self) -> Result<(), VolumeError> {
let cpc_path = self.file_name(".cpc");
if Path::new(&cpc_path).exists() {
// Marker present: the swap was decided. Finish whichever rename is
// still pending -- a crash may have completed only the .dat rename,
// so the lone remaining .cpx must still be applied, not abandoned.
// If neither temp file remains, finish_compact_swap clears the marker.
tracing::info!(
volume_id = self.id.0,
"rolling forward interrupted compaction commit"
);
return self.finish_compact_swap();
}
// No marker: roll back any orphan compaction temp files.
let mut rolled_back = false;
for ext in &[".cpd", ".cpx"] {
let p = self.file_name(ext);
if Path::new(&p).exists() {
tracing::info!(
volume_id = self.id.0,
"rolling back orphan compaction file {}",
ext
);
match fs::remove_file(&p) {
Ok(()) => {}
Err(e) if e.kind() == io::ErrorKind::NotFound => {}
Err(e) => return Err(e.into()),
}
rolled_back = true;
}
}
if rolled_back {
let _ = fs::remove_dir_all(self.file_name(".ldb"));
let _ = fs::remove_file(self.file_name(".rdb"));
}
Ok(())
}
/// Clean up leftover compaction files (.cpd, .cpx).
pub fn cleanup_compact(&self) -> Result<(), VolumeError> {
// An in-flight copy is still writing .cpd/.cpx.
if self.is_compacting() {
return Err(self.compacting_error());
}
// Refuse to unlink .cpd/.cpx while a .cpc marker exists: those temp files
// are the only inputs reconcile can roll forward to, so removing them
// mid-commit would strand a decided swap.
if Path::new(&self.file_name(".cpc")).exists() {
return Err(VolumeError::Io(io::Error::other(format!(
"volume {}: refusing cleanup while commit marker present",
self.id
))));
}
let cpd_path = self.file_name(".cpd");
let cpx_path = self.file_name(".cpx");
let cpldb_path = self.file_name(".cpldb");
let cpc_path = self.file_name(".cpc");
let e1 = fs::remove_file(&cpd_path);
let e2 = fs::remove_file(&cpx_path);
let e3 = fs::remove_dir_all(&cpldb_path);
let e4 = fs::remove_file(&cpc_path);
// Ignore NotFound errors
for e in [e1, e2, e3, e4] {
if let Err(e) = e
&& e.kind() != io::ErrorKind::NotFound
{
return Err(e.into());
}
}
Ok(())
}
/// Read any new needles appended during compaction and append them to .cpd/.cpx
fn makeup_diff(&mut self) -> Result<(), VolumeError> {
let old_idx_size = self.nm.as_ref().map_or(0, |nm| nm.index_file_size());
if old_idx_size == 0 || old_idx_size <= self.last_compact_index_offset {
return Ok(());
}
let old_super_block = &self.super_block;
if old_super_block.compaction_revision != self.last_compact_revision {
return Err(VolumeError::Io(io::Error::other(format!(
"current old dat file's compact revision {} is not the expected one {}",
old_super_block.compaction_revision, self.last_compact_revision
))));
}
// Read the new .cpd file's super block and verify its compaction revision is old + 1
let cpd_path_check = self.file_name(".cpd");
let mut cpd_file_check = File::open(&cpd_path_check)?;
let mut sb_buf = [0u8; SUPER_BLOCK_SIZE];
cpd_file_check.read_exact(&mut sb_buf)?;
let new_super_block = SuperBlock::from_bytes(&sb_buf)?;
let old_compact_revision = old_super_block.compaction_revision;
let new_compact_revision = new_super_block.compaction_revision;
if old_compact_revision + 1 != new_compact_revision {
return Err(VolumeError::Io(io::Error::other(format!(
"old dat file's compact revision {} + 1 does not equal new dat file's compact revision {}",
old_compact_revision, new_compact_revision
))));
}
let old_idx_path = self.file_name(".idx");
let mut old_idx_file = File::open(&old_idx_path)?;
// Read new entries from .idx
let mut incremented_entries = std::collections::HashMap::new();
let offset = self.last_compact_index_offset;
old_idx_file.seek(SeekFrom::Start(offset))?;
let entry_count = (old_idx_size - offset) / NEEDLE_MAP_ENTRY_SIZE as u64;
for _ in 0..entry_count {
let mut buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
old_idx_file.read_exact(&mut buf)?;
let (key, needle_offset, size) = crate::storage::types::idx_entry_from_bytes(&buf);
incremented_entries.insert(key, (needle_offset, size));
}
if incremented_entries.is_empty() {
return Ok(());
}
let cpd_path = self.file_name(".cpd");
let cpx_path = self.file_name(".cpx");
let mut dst_dat = OpenOptions::new().read(true).write(true).open(&cpd_path)?;
let mut dst_idx = OpenOptions::new().append(true).open(&cpx_path)?;
let mut dat_offset = dst_dat.seek(SeekFrom::End(0))?;
let padding_rem = dat_offset % NEEDLE_PADDING_SIZE as u64;
if padding_rem != 0 {
dat_offset += NEEDLE_PADDING_SIZE as u64 - padding_rem;
dst_dat.seek(SeekFrom::Start(dat_offset))?;
}
let version = self.version();
let old_dat_path = self.file_name(".dat");
let old_dat_file = File::open(&old_dat_path)?;
for (key, (needle_offset, size)) in incremented_entries {
if !needle_offset.is_zero() && !size.is_deleted() && size.0 > 0 {
let actual_size = crate::storage::needle::needle::get_actual_size(size, version);
let mut blob = vec![0u8; actual_size as usize];
read_exact_at(
&old_dat_file,
&mut blob,
needle_offset.to_actual_offset() as u64,
)?;
dst_dat.write_all(&blob)?;
let mut idx_entry_buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
crate::storage::types::idx_entry_to_bytes(
&mut idx_entry_buf,
key,
Offset::from_actual_offset(dat_offset as i64),
size,
);
dst_idx.write_all(&idx_entry_buf)?;
dat_offset += actual_size as u64;
} else {
let mut fake_del_needle = Needle {
id: key,
cookie: Cookie(0x12345678),
append_at_ns: SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap()
.as_nanos() as u64,
..Needle::default()
};
let bytes = fake_del_needle.write_bytes(version);
dst_dat.write_all(&bytes)?;
// Record the tombstone's real .dat offset, like the normal delete path,
// so a deletion left at the .dat tail stays visible to the integrity
// check on reload. Offset 0 hid the trailing tombstone and falsely
// flipped the volume read-only.
let mut idx_entry_buf = [0u8; NEEDLE_MAP_ENTRY_SIZE];
crate::storage::types::idx_entry_to_bytes(
&mut idx_entry_buf,
key,
Offset::from_actual_offset(dat_offset as i64),
Size(crate::storage::types::TOMBSTONE_FILE_SIZE.into()),
);
dst_idx.write_all(&idx_entry_buf)?;
dat_offset += bytes.len() as u64;
}
}
dst_dat.sync_all()?;
dst_idx.sync_all()?;
Ok(())
}
// ---- Sync / Close ----
pub fn sync_to_disk(&mut self) -> io::Result<()> {
if let Some(ref dat_file) = self.dat_file {
dat_file.sync_all()?;
}
if let Some(ref nm) = self.nm {
nm.sync()?;
}
Ok(())
}
pub fn close(&mut self) {
let dat_synced = if let Some(ref dat_file) = self.dat_file {
dat_file.sync_all().is_ok()
} else {
true
};
self.dat_file = None;
self.remote_dat_file = None;
// When the .dat flushed, checkpoint the index so the next load starts
// from the recorded .idx size. When it did not, skip the checkpoint:
// a durable index pointing past an unflushed .dat tail would make the
// reload's max_needle_end check mark the volume read-only. Without the
// checkpoint, META_IDX_SIZE stays at the last successful one and the
// reload replays the uncheckpointed tail (redb flushes on drop).
if let Some(ref mut nm) = self.nm {
if dat_synced {
nm.close();
} else {
tracing::warn!(
"volume {}: .dat sync failed on close, skipping index checkpoint",
self.id.0
);
nm.close_without_checkpoint();
}
}
self.nm = None;
}
/// Move this volume's `.idx` (and `.sdx` when present) from its current
/// index directory into `new_idx_dir`, then reload in place. Mirrors Go's
/// `Volume::RelocateIndexTo`: it is a no-op when the index is already there
/// or nothing is co-located to move, and is used to pull an index a decode
/// or reconstruct left beside the data back into the configured `-dir.idx`.
pub fn relocate_index_to(&mut self, new_idx_dir: &str) -> Result<(), VolumeError> {
// Moving .idx would strand the .cpx an in-flight copy writes beside it.
if self.is_compacting() {
return Err(self.compacting_error());
}
let _guard = self.data_file_access_control.write_lock();
if self.dir_idx == new_idx_dir {
return Ok(());
}
let old_base = volume_file_name(&self.dir_idx, &self.collection, self.id);
let old_idx = format!("{old_base}.idx");
if !Path::new(&old_idx).exists() {
return Ok(()); // nothing co-located to move
}
let new_base = volume_file_name(new_idx_dir, &self.collection, self.id);
let new_idx = format!("{new_base}.idx");
// Close the index and data handles before moving the index file.
let version = self.version();
self.close();
if let Err(e) = rename_or_copy(&old_idx, &new_idx) {
// Reopen against the old dir so the volume is not left down; surface
// a failed reopen since it leaves the volume unusable until reload.
if let Err(reopen) = self.load(true, false, 0, version) {
error!(
"relocate volume {}: reopen after failed .idx move: {}",
self.id, reopen
);
}
return Err(VolumeError::Io(io::Error::other(format!(
"relocate index for volume {}: move .idx: {e}",
self.id
))));
}
// The .sdx is a derived sorted index; move it when present, but a
// failure is not fatal — drop the stale copy so a reload rebuilds it.
let old_sdx = format!("{old_base}.sdx");
if Path::new(&old_sdx).exists() {
let new_sdx = format!("{new_base}.sdx");
if let Err(e) = rename_or_copy(&old_sdx, &new_sdx) {
warn!(
"relocate volume {}: move .sdx: {} (will rebuild)",
self.id, e
);
let _ = fs::remove_file(&old_sdx);
}
}
self.dir_idx = new_idx_dir.to_string();
self.load(true, false, 0, version)
}
/// Remove all volume files from disk.
/// Destroy removes everything related to this volume. When keep_remote_data
/// is true the cloud-tier object backing the volume is left intact — used
/// by moves where another server is taking over the same .vif.
pub fn destroy(
&mut self,
only_empty: bool,
only_garbage: bool,
keep_remote_data: bool,
) -> Result<(), VolumeError> {
// Either enabled check may pass: a volume with no live data qualifies
// whether it reads empty or as all garbage. Byte counters, not counts:
// index rows and live tallies drift apart on reload.
let empty_ok = only_empty && self.file_count() == 0;
let garbage_ok =
only_garbage && self.content_size() > 0 && self.deleted_size() >= self.content_size();
if (only_empty || only_garbage) && !empty_ok && !garbage_ok {
return Err(VolumeError::NotEmpty);
}
if self.is_compacting() {
return Err(self.compacting_error());
}
let (storage_name, storage_key) = self.remote_storage_name_key();
if !keep_remote_data
&& self.has_remote_file()
&& !storage_name.is_empty()
&& !storage_key.is_empty()
{
let backend = crate::remote_storage::s3_tier::global_s3_tier_registry()
.read()
.unwrap()
.get(&storage_name);
if let Some(backend) = backend {
if let Err(e) = backend.delete_file_blocking(&storage_key) {
warn!(
volume_id = self.id.0,
storage_name,
storage_key,
error = %e,
"failed to delete remote tier file during destroy"
);
}
} else {
warn!(
volume_id = self.id.0,
storage_name, storage_key, "remote tier backend not found during destroy"
);
}
}
// A regular volume and an EC volume for the same id share <base>.vif.
// When EC artefacts coexist on this disk (e.g. shards distributed onto
// a source replica before it is deleted), keep the .vif so removing the
// regular volume does not strip the EC volume's info file.
let keep_vif = self.shares_vif_with_ec_volume();
self.close();
remove_volume_files(&self.data_file_name(), keep_vif);
remove_volume_files(&self.index_file_name(), keep_vif);
Ok(())
}
/// Reports whether an EC volume for this id has a sealed .ecx on the same
/// disk, in which case its .vif is the same file as the regular volume's
/// and must outlive the regular volume's deletion. Mirrors the on-disk
/// portion of Go's Volume.sharesVifWithEcVolume / HasEcxFileOnDisk.
fn shares_vif_with_ec_volume(&self) -> bool {
let has_ecx = |base: &str| -> bool {
fs::metadata(format!("{}.ecx", base))
.map(|m| !m.is_dir() && m.len() > 0)
.unwrap_or(false)
};
if has_ecx(&volume_file_name(&self.dir_idx, &self.collection, self.id)) {
return true;
}
self.dir != self.dir_idx && has_ecx(&volume_file_name(&self.dir, &self.collection, self.id))
}
/// Matches Go's `checkReadWriteError` in `weed/storage/io_error.go`.
fn check_read_write_error(&self, err: Option<&io::Error>) {
self.io_errors.check_read_write_error(err);
}
pub fn get_io_error_state(&self) -> (Option<String>, i32, bool) {
self.io_errors.get_io_error_state()
}
pub fn should_quarantine(&self) -> bool {
self.io_errors.should_quarantine()
}
pub fn mark_io_quarantined(&self) {
self.io_errors.mark_io_quarantined();
}
pub fn reset_io_error_state(&self) {
self.io_errors.reset_io_error_state();
}
#[cfg(test)]
pub(crate) fn set_last_io_error_for_test(&self, err: Option<&str>) {
self.io_errors.set_last_io_error_for_test(err);
}
#[cfg(test)]
pub(crate) fn fail_next_fsync_for_test(&mut self, fail: bool) {
self.fail_fsync_for_test = fail;
}
/// Test-only read of the redb META `.idx` size through the live handle.
/// See `needle_map::test_support::live_meta_idx_size` for why durability
/// tests read this live instead of copying the open `.rdb` (Windows
/// mandatory file locking rejects reads of the locked file).
#[cfg(test)]
pub(crate) fn live_meta_idx_size_for_test(&self) -> Option<u64> {
match self.nm.as_ref() {
Some(NeedleMap::Redb(nm)) => nm.live_meta_idx_size(),
_ => None,
}
}
#[cfg(test)]
pub(crate) fn fail_next_idx_sync_for_test(&mut self, fail: bool) {
self.fail_idx_sync_for_test = fail;
}
#[cfg(test)]
pub(crate) fn fail_next_truncate_for_test(&mut self, fail: bool) {
self.fail_truncate_for_test = fail;
}
#[cfg(test)]
pub(crate) fn fail_append_for_test(&mut self, ids: &[NeedleId]) {
self.fail_append_for_test = ids.iter().copied().collect();
}
/// (.dat syncs, .idx syncs) attempted since the volume was opened.
#[cfg(test)]
pub(crate) fn sync_counts_for_test(&self) -> (usize, usize) {
(
self.dat_syncs_for_test.load(Ordering::Relaxed),
self.idx_syncs_for_test.load(Ordering::Relaxed),
)
}
#[cfg(test)]
pub(crate) fn set_last_modified_ts_for_test(&mut self, ts_seconds: u64) {
self.last_modified_ts_seconds = ts_seconds;
}
#[cfg(test)]
pub(crate) fn add_file_count_for_test(&self, delta: i64) {
if let Some(nm) = self.nm.as_ref() {
nm.add_file_count_for_test(delta);
}
}
}
// ============================================================================
// Helpers
// ============================================================================
/// A copy of a run-wide failure for each entry it fails.
fn run_error(e: &VolumeError) -> VolumeError {
match e {
VolumeError::Io(e) => VolumeError::Io(io::Error::new(e.kind(), e.to_string())),
e => VolumeError::Io(io::Error::other(e.to_string())),
}
}
/// Generate volume file base name: dir/collection_id or dir/id
/// Byte offset just past the needle's on-disk record. Mirrors Go's
/// needleDiskEnd.
pub(crate) fn needle_disk_end(offset: Offset, size: Size, version: Version) -> i64 {
offset.to_actual_offset() + get_actual_size(Size(size.raw() as i32), version)
}
fn size_mismatch_error(offset: i64, id: NeedleId, found: Size, expected: Size) -> VolumeError {
VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"entry not found: offset {} found id {:?} size {}, expected size {}",
offset, id, found.0, expected.0
),
))
}
/// Drop a partial row from the end of `.idx` and return the remaining length.
///
/// `.idx` is a run of fixed-size rows, and writers append at EOF: a short write
/// leaves a partial row behind, and every row appended after it lands off
/// alignment, so the next load parses the rest of the file as garbage. The
/// partial row is unrecoverable either way — every loader already skips it — so
/// removing it before anything can append is what keeps the file parseable.
/// Go instead refuses to load a volume whose `.idx` is not a whole number of
/// rows; trimming keeps it mountable, and the rows before the tear are intact.
fn trim_torn_idx_tail(idx_file: &File, idx_path: &str) -> Result<u64, VolumeError> {
let size = idx_file.metadata()?.len();
let aligned = size - size % NEEDLE_MAP_ENTRY_SIZE as u64;
if aligned != size {
warn!(
idx = %idx_path,
size, aligned, "dropping a partial row from the end of the index"
);
idx_file.set_len(aligned)?;
idx_file.sync_all()?;
}
Ok(aligned)
}
pub fn volume_file_name(dir: &str, collection: &str, id: VolumeId) -> String {
if collection.is_empty() {
format!("{}/{}", dir, id.0)
} else {
format!("{}/{}_{}", dir, collection, id.0)
}
}
/// Move `src` to `dst`, falling back to copy+remove when the two live on
/// different filesystems (a real `-dir.idx` split). Mirrors Go's
/// `RenameOrCopyFile`: any rename error other than cross-device is fatal, and a
/// failed copy or a source that cannot be removed rolls the copy back so a
/// failure never leaves two divergent files.
fn rename_or_copy(src: &str, dst: &str) -> io::Result<()> {
match fs::rename(src, dst) {
Ok(()) => return Ok(()),
Err(e) if e.raw_os_error() != Some(libc::EXDEV) => return Err(e),
Err(_) => {}
}
fs::copy(src, dst)?;
// fsync the destination before dropping the source, matching Go's out.Sync().
if let Err(e) = OpenOptions::new()
.write(true)
.open(dst)
.and_then(|f| f.sync_all())
{
let _ = fs::remove_file(dst);
return Err(e);
}
if let Err(e) = fs::remove_file(src) {
let _ = fs::remove_file(dst);
return Err(e);
}
Ok(())
}
/// Generate a monotonically increasing append timestamp.
fn get_append_at_ns(last: u64) -> u64 {
let now = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap_or_default()
.as_nanos() as u64;
if now <= last { last + 1 } else { now }
}
/// Remove all files associated with a volume.
/// .dat/.idx removals log at info level so destructive calls are traceable.
/// fsync the parent directory of `path` so a rename/create/unlink inside it is
/// durable, propagating a sync failure so the commit path can abort rather than
/// proceed with an undurable rename or marker. A path with no openable parent is
/// tolerated; directory fsync is unsupported on Windows, so it is a no-op there
/// (matching the Go fsyncDir helper, which skips it too).
pub(crate) fn fsync_dir(path: &str) -> io::Result<()> {
#[cfg(windows)]
{
let _ = path;
Ok(())
}
#[cfg(not(windows))]
{
if let Some(parent) = Path::new(path).parent()
&& let Ok(d) = File::open(parent)
{
return d.sync_all();
}
Ok(())
}
}
pub(crate) fn remove_volume_files(base: &str, keep_vif: bool) {
for ext in &[
".dat",
".idx",
".vif",
".sdx",
".cpd",
".cpx",
".cpc",
".note",
".rdb",
".unavailable",
] {
if *ext == ".vif" && keep_vif {
continue;
}
let path = format!("{}{}", base, ext);
let size = fs::metadata(&path).map(|m| m.len()).unwrap_or(0);
let existed = fs::remove_file(&path).is_ok();
if existed && (*ext == ".dat" || *ext == ".idx") {
tracing::info!("removed volume file {} (size={})", path, size);
}
}
// leveldb uses a directory
let _ = fs::remove_dir_all(format!("{}.ldb", base));
}
// ============================================================================
// ScanVolumeFile — iterate all needles in a .dat file
// ============================================================================
/// Callback for scanning needles in a volume file.
pub trait VolumeFileVisitor {
fn visit_super_block(&mut self, sb: &SuperBlock) -> Result<(), VolumeError>;
fn read_needle_body(&self) -> bool;
fn visit_needle(&mut self, n: &Needle, offset: i64) -> Result<(), VolumeError>;
}
/// Scan all needles in a volume's .dat file.
pub fn scan_volume_file(
dat_path: &str,
visitor: &mut dyn VolumeFileVisitor,
) -> Result<(), VolumeError> {
let mut file = File::open(dat_path)?;
// Read super block
let mut sb_buf = [0u8; SUPER_BLOCK_SIZE];
file.read_exact(&mut sb_buf)?;
let sb = SuperBlock::from_bytes(&sb_buf)?;
visitor.visit_super_block(&sb)?;
let version = sb.version;
let mut offset = sb.block_size() as i64;
loop {
// Read needle header
let mut header = [0u8; NEEDLE_HEADER_SIZE];
file.seek(SeekFrom::Start(offset as u64))?;
match file.read_exact(&mut header) {
Ok(()) => {}
Err(e) if e.kind() == io::ErrorKind::UnexpectedEof => break,
Err(e) => return Err(e.into()),
}
let (_cookie, _id, size) = Needle::parse_header(&header);
if size.0 == 0 && _id.is_empty() {
break; // end of valid data
}
// A negative size is a corrupt header, and body_length would advance the
// walk backwards from it. The index rebuild this feeds salvages the
// records before it, like Go's rebuild scanner stopping on io.EOF.
if size.0 < 0 {
break;
}
let body_length = needle::needle_body_length(size, version);
let total_size = NEEDLE_HEADER_SIZE as i64 + body_length;
// Skip full body parsing for deleted needles (tombstone or negative size)
if size.is_deleted() || size.0 <= 0 {
let mut n = Needle::default();
n.read_header(&header);
visitor.visit_needle(&n, offset)?;
} else if visitor.read_needle_body() {
let mut buf = vec![0u8; total_size as usize];
file.seek(SeekFrom::Start(offset as u64))?;
file.read_exact(&mut buf)?;
let mut n = Needle::default();
n.read_bytes(&buf, offset, size, version)?;
visitor.visit_needle(&n, offset)?;
} else {
let mut n = Needle::default();
n.read_header(&header);
visitor.visit_needle(&n, offset)?;
}
offset += total_size;
}
Ok(())
}
/// Reserve disk blocks for a file without changing its visible size.
/// On Linux, uses `fallocate(FALLOC_FL_KEEP_SIZE)` to actually reserve blocks.
/// On other platforms, this is a no-op.
fn preallocate_file(file: &File, size: u64) {
#[cfg(target_os = "linux")]
{
use std::os::unix::io::AsRawFd;
let fd = file.as_raw_fd();
// FALLOC_FL_KEEP_SIZE = 1: allocate blocks without changing file size.
//
// SAFETY: `fd` is borrowed from the live `&File` the caller owns, so
// it stays open for the call; the remaining arguments are plain
// scalars; and failure is reported in the return value, which is
// checked below before `last_os_error()` reads errno.
let ret = unsafe { libc::fallocate(fd, 1, 0, size as libc::off_t) };
if ret == 0 {
tracing::info!(bytes = size, "preallocated disk space");
} else {
tracing::warn!(
bytes = size,
error = %io::Error::last_os_error(),
"fallocate failed"
);
}
}
#[cfg(not(target_os = "linux"))]
{
let _ = (file, size);
tracing::debug!(bytes = size, "preallocation not supported on this platform");
}
}
// ============================================================================
// Tests
// ============================================================================
/// An OS error the platform reports for failing storage media, which is
/// what counts toward the I/O error streak.
#[cfg(test)]
fn media_error_for_test() -> io::Error {
#[cfg(unix)]
{
io::Error::from_raw_os_error(libc::EIO)
}
#[cfg(windows)]
{
const ERROR_IO_DEVICE: i32 = 1117;
io::Error::from_raw_os_error(ERROR_IO_DEVICE)
}
#[cfg(not(any(unix, windows)))]
{
io::Error::other("injected media error")
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::storage::needle::crc::CRC;
use tempfile::TempDir;
fn spawn_fake_s3_server(body: Vec<u8>) -> (String, tokio::sync::oneshot::Sender<()>) {
use axum::Router;
use axum::http::{HeaderMap, HeaderValue, StatusCode, header};
use axum::routing::any;
let body = Arc::new(body);
let listener = std::net::TcpListener::bind("127.0.0.1:0").unwrap();
let addr = listener.local_addr().unwrap();
listener.set_nonblocking(true).unwrap();
let (shutdown_tx, shutdown_rx) = tokio::sync::oneshot::channel::<()>();
let (ready_tx, ready_rx) = std::sync::mpsc::channel::<()>();
std::thread::spawn(move || {
let runtime = tokio::runtime::Builder::new_current_thread()
.enable_all()
.build()
.unwrap();
runtime.block_on(async move {
let app = Router::new().fallback(any(move |headers: HeaderMap| {
let body = body.clone();
async move {
let bytes = body.as_ref();
if let Some(range) = headers
.get(header::RANGE)
.and_then(|value| value.to_str().ok())
&& let Some(spec) = range.strip_prefix("bytes=")
{
let (start, end) = spec.split_once('-').unwrap();
let start = start.parse::<usize>().unwrap();
let end = if end.is_empty() {
bytes.len().saturating_sub(1)
} else {
end.parse::<usize>().unwrap()
}
.min(bytes.len().saturating_sub(1));
let chunk = bytes[start..=end].to_vec();
let mut response_headers = HeaderMap::new();
response_headers.insert(
header::CONTENT_LENGTH,
HeaderValue::from_str(&chunk.len().to_string()).unwrap(),
);
response_headers.insert(
header::CONTENT_RANGE,
HeaderValue::from_str(&format!(
"bytes {}-{}/{}",
start,
end,
bytes.len()
))
.unwrap(),
);
return (StatusCode::PARTIAL_CONTENT, response_headers, chunk);
}
let mut response_headers = HeaderMap::new();
response_headers.insert(
header::CONTENT_LENGTH,
HeaderValue::from_str(&bytes.len().to_string()).unwrap(),
);
(StatusCode::OK, response_headers, bytes.to_vec())
}
}));
let listener = tokio::net::TcpListener::from_std(listener).unwrap();
let _ = ready_tx.send(());
axum::serve(listener, app)
.with_graceful_shutdown(async move {
let _ = shutdown_rx.await;
})
.await
.unwrap();
});
});
// Wait for the server thread to be ready before returning.
ready_rx.recv().unwrap();
(format!("http://{}", addr), shutdown_tx)
}
fn make_test_volume(dir: &str) -> Volume {
Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap()
}
fn write_test_needle(v: &mut Volume, id: u64, data: &[u8]) -> u64 {
let mut n = Needle {
id: NeedleId(id),
cookie: Cookie(0x12345678),
data: data.to_vec(),
data_size: data.len() as u32,
flags: 0,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
n.append_at_ns
}
/// Run a plan to completion, keeping owned copies of every record.
fn scan_all(plan: &DatScanPlan) -> Vec<(Vec<u8>, Vec<u8>, u64)> {
let mut records = Vec::new();
plan.scan(|needle| {
records.push((
needle.header.to_vec(),
needle.body.to_vec(),
needle.append_at_ns,
));
ControlFlow::Continue(())
})
.unwrap();
records
}
#[test]
fn test_data_file_access_control_blocks_writer_until_reader_releases() {
let control = Arc::new(DataFileAccessControl::default());
let read_lease = control.read_lock();
let writer_control = control.clone();
let acquired = Arc::new(std::sync::atomic::AtomicBool::new(false));
let acquired_clone = acquired.clone();
let writer = std::thread::spawn(move || {
let _write_lease = writer_control.write_lock();
acquired_clone.store(true, std::sync::atomic::Ordering::Relaxed);
});
std::thread::sleep(std::time::Duration::from_millis(50));
assert!(!acquired.load(std::sync::atomic::Ordering::Relaxed));
drop(read_lease);
writer.join().unwrap();
assert!(acquired.load(std::sync::atomic::Ordering::Relaxed));
}
#[test]
fn test_volume_file_name() {
assert_eq!(volume_file_name("/data", "", VolumeId(1)), "/data/1");
assert_eq!(
volume_file_name("/data", "pics", VolumeId(42)),
"/data/pics_42"
);
}
#[test]
fn test_volume_create_and_load() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let v = make_test_volume(dir);
assert_eq!(v.version(), VERSION_3);
assert_eq!(v.file_count(), 0);
assert_eq!(v.content_size(), 0);
// .dat and .idx files should exist
assert!(Path::new(&v.file_name(".dat")).exists());
assert!(Path::new(&v.file_name(".idx")).exists());
}
#[test]
fn binary_search_reports_is_last_when_caller_is_caught_up() {
// is_last drives the caught-up heartbeat; if it stops reporting true
// for an up-to-date caller, the tail sender re-reads the whole volume.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut newest_ns = 0u64;
for id in 1..=3u64 {
let mut n = Needle {
id: NeedleId(id),
cookie: Cookie(0x12345678),
data: vec![b'x'; 64],
data_size: 64,
flags: 0,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
newest_ns = n.append_at_ns;
}
assert!(newest_ns > 0, "needles must carry an append timestamp");
// Caught up: asking from the newest timestamp has nothing newer.
let (_offset, is_last) = v.binary_search_by_append_at_ns(newest_ns).unwrap();
assert!(
is_last,
"a caller at the newest append_at_ns must be reported as caught up"
);
// And from beyond the newest, likewise.
let (_offset, is_last_future) = v
.binary_search_by_append_at_ns(newest_ns + 1_000_000_000)
.unwrap();
assert!(
is_last_future,
"a caller ahead of the newest needle must be reported as caught up"
);
}
#[test]
fn binary_search_does_not_report_is_last_when_data_is_newer() {
// Complement: a behind caller must NOT be caught up, or tailing loses data.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut first_ns = 0u64;
for id in 1..=3u64 {
let mut n = Needle {
id: NeedleId(id),
cookie: Cookie(0x12345678),
data: vec![b'y'; 64],
data_size: 64,
flags: 0,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
if id == 1 {
first_ns = n.append_at_ns;
}
}
let (_offset, is_last) = v.binary_search_by_append_at_ns(first_ns).unwrap();
assert!(
!is_last,
"a caller behind the newest needle must NOT be reported as caught up"
);
}
#[test]
fn binary_search_does_not_report_is_last_when_only_a_delete_is_newer() {
// A delete is data the tail must ship. A caught-up caller before the
// delete must be sent to scan from the tombstone, not handed a heartbeat.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut newest_ns = 0u64;
for id in 1..=3u64 {
let mut n = Needle {
id: NeedleId(id),
cookie: Cookie(0x12345678),
data: vec![b'z'; 64],
data_size: 64,
flags: 0,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
newest_ns = n.append_at_ns;
}
let (_offset, is_last) = v.binary_search_by_append_at_ns(newest_ns).unwrap();
assert!(is_last, "precondition: the caller starts caught up");
v.delete_needle(&mut Needle {
id: NeedleId(2),
cookie: Cookie(0x12345678),
..Needle::default()
})
.unwrap();
let (offset, is_last) = v.binary_search_by_append_at_ns(newest_ns).unwrap();
assert!(
!is_last,
"a caller older than a trailing delete must NOT be reported as caught up"
);
// Scan and filter exactly as volume_tail_sender does: the tombstone,
// and only the tombstone, is newer than the caller.
let shipped: Vec<u64> =
scan_all(&v.dat_scan_plan(offset.to_actual_offset() as u64).unwrap())
.into_iter()
.map(|(_, _, append_at_ns)| append_at_ns)
.filter(|&append_at_ns| append_at_ns > newest_ns)
.collect();
assert_eq!(shipped.len(), 1, "the tail must ship the tombstone");
}
#[test]
fn binary_search_on_compacted_volume_still_reports_a_later_write() {
// Compaction rewrites .idx in needle-id order, so append_at_ns is no
// longer monotonic by row: an overwritten key can sit before an older
// one. The caller's since_ns is the last row's timestamp, so such rows
// were already in the files it copied. A write made afterwards is
// appended as the final row, which the search cannot step past.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let write = |v: &mut Volume, id: u64, byte: u8| {
let mut n = Needle {
id: NeedleId(id),
cookie: Cookie(0x12345678),
data: vec![byte; 64],
data_size: 64,
flags: 0,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
n.append_at_ns
};
write(&mut v, 1, b'c');
let key2_ns = write(&mut v, 2, b'c');
let key1_ns = write(&mut v, 1, b'd'); // genuine overwrite: different data
assert!(
key1_ns > key2_ns,
"precondition: the overwrite must append a newer record, not dedup"
);
v.compact_by_index(0, 0, |_| true).unwrap();
v.commit_compact().unwrap();
let (_offset, is_last) = v.binary_search_by_append_at_ns(key2_ns).unwrap();
assert!(
is_last,
"precondition: compaction ordered .idx by key, so key 2 is the final row"
);
let key3_ns = write(&mut v, 3, b'c');
let (offset, is_last) = v.binary_search_by_append_at_ns(key2_ns).unwrap();
assert!(
!is_last,
"a write after compaction is the final row and must not be hidden"
);
let shipped: Vec<u64> =
scan_all(&v.dat_scan_plan(offset.to_actual_offset() as u64).unwrap())
.into_iter()
.map(|(_, _, append_at_ns)| append_at_ns)
.filter(|&append_at_ns| append_at_ns > key2_ns)
.collect();
assert!(
shipped.contains(&key3_ns),
"the tail must ship the write made after compaction"
);
}
#[test]
fn dat_scan_plan_visits_needles_and_tombstones_in_file_order() {
// The tail ships exactly what the scan visits, so the scan must cover
// the file record by record: the on-disk bytes, in file order,
// tombstones included.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let written: Vec<u64> = (1..=3u64)
.map(|id| write_test_needle(&mut v, id, b"scan"))
.collect();
v.delete_needle(&mut Needle {
id: NeedleId(2),
cookie: Cookie(0x12345678),
..Needle::default()
})
.unwrap();
let sb_size = v.super_block.block_size() as u64;
let records = scan_all(&v.dat_scan_plan(sb_size).unwrap());
let dat = std::fs::read(v.file_name(".dat")).unwrap();
assert_eq!(records.len(), 4, "three writes and one tombstone");
let mut offset = sb_size as usize;
for (header, body, _) in &records {
assert_eq!(&dat[offset..offset + header.len()], header.as_slice());
offset += header.len();
assert_eq!(&dat[offset..offset + body.len()], body.as_slice());
offset += body.len();
}
assert_eq!(offset, dat.len(), "the scan must reach the end of the file");
let ids: Vec<NeedleId> = records
.iter()
.map(|(header, _, _)| Needle::parse_header(header).1)
.collect();
assert_eq!(
ids,
vec![NeedleId(1), NeedleId(2), NeedleId(3), NeedleId(2)]
);
let stamps: Vec<u64> = records.iter().map(|r| r.2).collect();
assert_eq!(&stamps[..3], written.as_slice());
assert!(stamps[3] > stamps[2], "the tombstone is the newest record");
}
#[cfg(unix)]
#[test]
fn dat_scan_plan_outlives_a_vacuum_commit() {
// The tail sender drops the store guard before scanning, so a vacuum
// commit can land mid-scan. The commit renames .cpd over .dat; a plan
// taken before it must keep reading the file its offset was resolved
// against, not the compacted one.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, b"first");
write_test_needle(&mut v, 2, b"second");
write_test_needle(&mut v, 1, b"first, overwritten");
let sb_size = v.super_block.block_size() as u64;
let before = scan_all(&v.dat_scan_plan(sb_size).unwrap());
let plan = v.dat_scan_plan(sb_size).unwrap();
let old_len = std::fs::metadata(v.file_name(".dat")).unwrap().len();
v.compact_by_index(0, 0, |_| true).unwrap();
v.commit_compact().unwrap();
let new_len = std::fs::metadata(v.file_name(".dat")).unwrap().len();
assert!(
new_len < old_len,
"precondition: the commit swapped in a smaller .dat"
);
assert_eq!(scan_all(&plan), before);
}
#[cfg(unix)]
#[test]
fn dat_scan_plan_outlives_volume_destroy() {
// Unmount or delete can land mid-scan too. The plan holds a handle,
// not a path, so it finishes reading what was there.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, b"going");
write_test_needle(&mut v, 2, b"gone");
let sb_size = v.super_block.block_size() as u64;
let before = scan_all(&v.dat_scan_plan(sb_size).unwrap());
let plan = v.dat_scan_plan(sb_size).unwrap();
let dat_path = v.file_name(".dat");
v.destroy(false, false, false).unwrap();
assert!(
!Path::new(&dat_path).exists(),
"precondition: destroy removed .dat"
);
assert_eq!(scan_all(&plan), before);
}
#[test]
fn dat_scan_plan_fails_when_the_snapshot_is_truncated() {
// Unmount followed by a VolumeCopy of the same id reopens .dat with
// truncate(true) on the inode a running plan holds. Every byte below
// the captured end existed when the plan was taken, so a short read
// there is an incomplete pass, not the end of the data: the stream
// must fail rather than let volume.move take a partial tail as
// complete.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, b"kept");
write_test_needle(&mut v, 2, b"cut off");
let sb_size = v.super_block.block_size() as u64;
let first_record_len = {
let records = scan_all(&v.dat_scan_plan(sb_size).unwrap());
assert_eq!(records.len(), 2, "precondition: two records on disk");
(records[0].0.len() + records[0].1.len()) as u64
};
let plan = v.dat_scan_plan(sb_size).unwrap();
let end = plan.end;
// Truncate through the header of the second record.
let second = sb_size + first_record_len;
OpenOptions::new()
.write(true)
.open(v.file_name(".dat"))
.unwrap()
.set_len(second + NEEDLE_HEADER_SIZE as u64 / 2)
.unwrap();
let mut visited = Vec::new();
let err = plan
.scan(|n| {
visited.push(n.append_at_ns);
ControlFlow::Continue(())
})
.unwrap_err();
assert_eq!(visited.len(), 1, "the intact first record is still visited");
match err {
VolumeError::Io(e) => {
assert_eq!(e.kind(), io::ErrorKind::UnexpectedEof);
let msg = e.to_string();
assert!(msg.contains("truncated"), "{msg}");
assert!(msg.contains(&format!("snapshot end {end}")), "{msg}");
}
other => panic!("expected an I/O error, got {other:?}"),
}
}
#[test]
fn dat_scan_plan_stops_at_the_snapshot_end() {
// The end bound is read under the guard, while no append can be in
// flight. A record appended after the plan belongs to the next pass.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let first_ns = write_test_needle(&mut v, 1, b"before");
let plan = v.dat_scan_plan(v.super_block.block_size() as u64).unwrap();
write_test_needle(&mut v, 2, b"after");
let stamps: Vec<u64> = scan_all(&plan).into_iter().map(|r| r.2).collect();
assert_eq!(stamps, vec![first_ns]);
}
#[test]
fn dat_scan_plan_stops_when_the_visitor_breaks() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for id in 1..=3u64 {
write_test_needle(&mut v, id, b"brk");
}
let mut visited = 0;
v.dat_scan_plan(v.super_block.block_size() as u64)
.unwrap()
.scan(|_| {
visited += 1;
ControlFlow::Break(())
})
.unwrap();
assert_eq!(visited, 1);
}
#[test]
fn dat_scan_plan_ends_the_pass_at_a_corrupt_header() {
// A negative size is a corrupt header: the record length it implies
// cannot advance the scan past it, so the pass fails. A size running
// past the end bound, like one near i32::MAX, cannot be a complete
// record and ends the pass after the records before it, without
// reading or allocating the bogus body.
for bad_size in [-1000, i32::MAX] {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, b"good");
let sb_size = v.super_block.block_size() as u64;
let dat_path = v.file_name(".dat");
let start = sb_size as usize;
let mut bad =
std::fs::read(&dat_path).unwrap()[start..start + NEEDLE_HEADER_SIZE].to_vec();
Size(bad_size).to_bytes(&mut bad[COOKIE_SIZE + NEEDLE_ID_SIZE..NEEDLE_HEADER_SIZE]);
// Bytes after the header, so the size check stops the walk, not EOF.
bad.extend_from_slice(&[0u8; 64]);
OpenOptions::new()
.append(true)
.open(&dat_path)
.unwrap()
.write_all(&bad)
.unwrap();
let mut visited = 0;
let result = v.dat_scan_plan(sb_size).unwrap().scan(|_| {
visited += 1;
ControlFlow::Continue(())
});
if bad_size < 0 {
let err = result.unwrap_err();
assert!(
matches!(&err, VolumeError::Io(e) if e.kind() == io::ErrorKind::InvalidData),
"size {bad_size}: expected a corrupt-data error, got {err:?}"
);
} else {
result.unwrap();
}
assert_eq!(visited, 1, "size {bad_size}: only the good record");
}
}
#[test]
fn dat_scan_plan_fails_when_the_record_body_is_corrupt() {
// A record can fit inside the captured snapshot and still fail needle
// parsing. The tail must surface that failure instead of emitting a
// raw record with a zero append timestamp and reporting a clean pass.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, b"good");
let sb_size = v.super_block.block_size() as u64;
let plan = v.dat_scan_plan(sb_size).unwrap();
let mut dat = OpenOptions::new()
.write(true)
.open(v.file_name(".dat"))
.unwrap();
dat.seek(SeekFrom::Start(sb_size + NEEDLE_HEADER_SIZE as u64))
.unwrap();
dat.write_all(&u32::MAX.to_be_bytes()).unwrap();
drop(dat);
let mut visited = 0;
let err = plan
.scan(|_| {
visited += 1;
ControlFlow::Continue(())
})
.unwrap_err();
assert_eq!(visited, 0, "the corrupt record must not be emitted");
assert!(
matches!(err, VolumeError::Needle(_)),
"expected a needle parse error, got {err:?}"
);
}
fn append_corrupt_header(v: &Volume, size: i32) {
let mut dat = OpenOptions::new()
.append(true)
.open(v.file_name(".dat"))
.unwrap();
let mut corrupt = [0u8; NEEDLE_HEADER_SIZE];
NeedleId(99).to_bytes(&mut corrupt[4..12]);
Size(size).to_bytes(&mut corrupt[12..16]);
dat.write_all(&corrupt).unwrap();
}
#[test]
fn dat_scan_plan_fails_at_a_header_it_cannot_advance_past() {
// A size so negative the record length is zero or less cannot be
// advanced past; the pass must fail rather than end early and let a
// truncated tail look complete.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, b"good");
append_corrupt_header(&v, -100);
let plan = v.dat_scan_plan(v.super_block.block_size() as u64).unwrap();
let mut visited = 0;
let err = plan
.scan(|_| {
visited += 1;
ControlFlow::Continue(())
})
.unwrap_err();
assert_eq!(
visited, 1,
"the record before the corrupt header is visited"
);
assert!(
matches!(&err, VolumeError::Io(e) if e.kind() == io::ErrorKind::InvalidData),
"expected a corrupt-data error, got {err:?}"
);
}
#[test]
fn test_volume_write_read() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
// Write a needle
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0x12345678),
data: b"hello world".to_vec(),
data_size: 11,
flags: 0,
..Needle::default()
};
let (offset, size, unchanged) = v.write_needle(&mut n, true, false).unwrap();
assert!(!unchanged);
assert!(offset > 0); // after superblock
assert!(size.0 > 0);
assert_eq!(v.file_count(), 1);
// Read it back
let mut read_n = Needle {
id: NeedleId(1),
..Needle::default()
};
let count = v.read_needle(&mut read_n).unwrap();
assert_eq!(count, 11);
assert_eq!(read_n.data, b"hello world");
assert_eq!(read_n.cookie, Cookie(0x12345678));
}
/// A durable write goes down the same path and lands the same data; the
/// flush is extra work, not different work.
#[test]
fn test_volume_write_needle_fsync() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0x12345678),
data: b"durable payload".to_vec(),
data_size: 15,
flags: 0,
..Needle::default()
};
let (offset, _, unchanged) = v.write_needle(&mut n, true, true).unwrap();
assert!(!unchanged);
assert!(offset > 0);
let mut read_n = Needle {
id: NeedleId(1),
..Needle::default()
};
assert_eq!(v.read_needle(&mut read_n).unwrap(), 15);
assert_eq!(read_n.data, b"durable payload");
// rewriting the same content still dedups, so there is nothing to flush
let mut same = Needle {
id: NeedleId(1),
cookie: Cookie(0x12345678),
data: b"durable payload".to_vec(),
data_size: 15,
flags: 0,
..Needle::default()
};
let (_, _, unchanged) = v.write_needle(&mut same, true, true).unwrap();
assert!(unchanged);
}
/// A durable write whose flush fails must leave nothing behind: the append
/// comes off the .dat and the mapping it would have replaced still stands,
/// with the volume's own accounting untouched.
#[test]
fn test_write_needle_failed_fsync_keeps_prior_mapping() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut kept = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"first-copy".to_vec(),
data_size: 10,
..Needle::default()
};
v.write_needle(&mut kept, true, true).unwrap();
let prior = v.nm.as_ref().unwrap().get(NeedleId(1)).unwrap().unwrap();
let dat_len_before = std::fs::metadata(v.file_name(".dat")).unwrap().len();
let file_count_before = v.file_count();
let content_size_before = v.content_size();
v.fail_next_fsync_for_test(true);
let mut replacement = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"second-copy".to_vec(),
data_size: 11,
..Needle::default()
};
v.write_needle(&mut replacement, true, true).unwrap_err();
v.fail_next_fsync_for_test(false);
assert_eq!(
std::fs::metadata(v.file_name(".dat")).unwrap().len(),
dat_len_before,
"the unflushed append should be off the .dat"
);
let now = v.nm.as_ref().unwrap().get(NeedleId(1)).unwrap().unwrap();
assert_eq!(
now.offset, prior.offset,
"the mapping should never have moved"
);
assert_eq!(now.size, prior.size);
assert_eq!(
v.file_count(),
file_count_before,
"a rejected write must not count towards the volume"
);
assert_eq!(v.content_size(), content_size_before);
let mut read_n = Needle {
id: NeedleId(1),
..Needle::default()
};
assert_eq!(v.read_needle(&mut read_n).unwrap(), 10);
assert_eq!(read_n.data, b"first-copy");
}
/// A durable write that dedups against an earlier non-durable one still has
/// to flush the index. The row that earlier write left behind can be sitting
/// in the page cache, so acking without flushing is the same false promise
/// as skipping the flush on a fresh append.
#[test]
fn test_dedup_durable_write_flushes_the_index() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut first = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"same-content".to_vec(),
data_size: 12,
..Needle::default()
};
v.write_needle(&mut first, true, false).unwrap();
v.fail_next_idx_sync_for_test(true);
let mut same = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"same-content".to_vec(),
data_size: 12,
..Needle::default()
};
let result = v.write_needle(&mut same, true, true);
v.fail_next_idx_sync_for_test(false);
assert!(
result.is_err(),
"a dedup hit must not be acked without flushing the index"
);
assert!(
v.is_read_only(),
"and the failed index flush quarantines the volume, dedup or not"
);
}
/// A durable write whose index flush fails leaves the .dat record down and
/// its row published but not guaranteed to come back. The volume stops
/// taking writes rather than appending more past a record whose index may
/// not survive a restart.
#[test]
fn test_write_needle_failed_idx_sync_quarantines_volume() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
v.fail_next_idx_sync_for_test(true);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"index-never-flushed".to_vec(),
data_size: 19,
..Needle::default()
};
v.write_needle(&mut n, true, true).unwrap_err();
v.fail_next_idx_sync_for_test(false);
assert!(
v.is_read_only(),
"a volume whose index flush failed must stop taking writes"
);
let mut later = Needle {
id: NeedleId(2),
cookie: Cookie(0xbb),
data: b"should-be-refused".to_vec(),
data_size: 17,
..Needle::default()
};
assert!(
matches!(
v.write_needle(&mut later, true, false),
Err(VolumeError::ReadOnly(_))
),
"later writes must not append past the record whose index is in doubt"
);
}
/// A failed fsync whose .dat rollback cannot complete leaves the tail
/// unverified, so the volume fails closed: reads and writes are refused,
/// the state is persisted, and a reload stays quarantined.
#[test]
fn test_failed_rollback_marks_volume_unavailable() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
v.fail_next_fsync_for_test(true);
v.fail_next_truncate_for_test(true);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"never-landed".to_vec(),
data_size: 12,
..Needle::default()
};
v.write_needle(&mut n, true, true).unwrap_err();
v.fail_next_fsync_for_test(false);
v.fail_next_truncate_for_test(false);
assert!(v.is_read_only());
assert!(v.unavailable_error().is_some());
assert!(
v.should_quarantine(),
"an unavailable volume must stay out of heartbeats"
);
assert!(Path::new(&v.file_name(".unavailable")).exists());
let mut read_n = Needle {
id: NeedleId(1),
..Needle::default()
};
assert!(matches!(
v.read_needle(&mut read_n),
Err(VolumeError::Unavailable(_))
));
let mut later = Needle {
id: NeedleId(2),
cookie: Cookie(0xbb),
data: b"should-be-refused".to_vec(),
data_size: 17,
..Needle::default()
};
assert!(matches!(
v.write_needle(&mut later, true, true),
Err(VolumeError::Unavailable(_))
));
assert!(
v.set_writable().is_err(),
"a manual writable transition must not bypass quarantine"
);
assert!(v.unavailable_error().is_some());
let reloaded = reload_volume(dir);
assert!(reloaded.unavailable_error().is_some());
assert!(reloaded.is_read_only());
assert!(reloaded.should_quarantine());
let mut read_n = Needle {
id: NeedleId(1),
..Needle::default()
};
assert!(reloaded.read_needle(&mut read_n).is_err());
}
/// Same for a needle the failed write introduced: it never becomes visible,
/// rather than being published and then tombstoned back out.
#[test]
fn test_write_needle_failed_fsync_publishes_nothing() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let dat_len_before = std::fs::metadata(v.file_name(".dat")).unwrap().len();
let file_count_before = v.file_count();
v.fail_next_fsync_for_test(true);
let mut fresh = Needle {
id: NeedleId(7),
cookie: Cookie(0x77),
data: b"never-landed".to_vec(),
data_size: 12,
..Needle::default()
};
v.write_needle(&mut fresh, true, true).unwrap_err();
v.fail_next_fsync_for_test(false);
assert_eq!(
std::fs::metadata(v.file_name(".dat")).unwrap().len(),
dat_len_before
);
assert!(
v.nm.as_ref().unwrap().get(NeedleId(7)).unwrap().is_none(),
"a needle that never reached the disk must not be indexed at all"
);
assert_eq!(v.file_count(), file_count_before);
assert_eq!(
v.deleted_count(),
0,
"nothing was written, so nothing was deleted"
);
let mut read_n = Needle {
id: NeedleId(7),
..Needle::default()
};
assert!(v.read_needle(&mut read_n).is_err());
}
fn batch_needle(id: u64, cookie: u32, data: &[u8]) -> Needle {
Needle {
id: NeedleId(id),
cookie: Cookie(cookie),
data: data.to_vec(),
data_size: data.len() as u32,
..Needle::default()
}
}
fn dat_len(v: &Volume) -> u64 {
std::fs::metadata(v.file_name(".dat")).unwrap().len()
}
/// A batch of durable writes shares one .dat sync and one .idx sync,
/// where writing them one at a time syncs both files per needle.
#[test]
fn test_grouped_fsync_writes_share_one_sync() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut writes: Vec<_> = (1..=10u64)
.map(|id| {
(
batch_needle(id, 0xaa, format!("body-{id}").as_bytes()),
true,
)
})
.collect();
let results = v.write_needles_grouped(&mut writes);
assert!(results.iter().all(|r| matches!(r, Ok((_, _, false)))));
assert_eq!(v.sync_counts_for_test(), (1, 1));
for id in 1..=10u64 {
let mut n = Needle {
id: NeedleId(id),
..Needle::default()
};
v.read_needle(&mut n).unwrap();
assert_eq!(n.data, format!("body-{id}").as_bytes());
}
// A batch with no durable write syncs nothing, as before.
let mut writes: Vec<_> = (11..=20u64)
.map(|id| (batch_needle(id, 0xaa, b"lazy"), false))
.collect();
let results = v.write_needles_grouped(&mut writes);
assert!(results.iter().all(|r| r.is_ok()));
assert_eq!(v.sync_counts_for_test(), (1, 1));
// The per-needle path pays both syncs for every durable write.
let tmp2 = TempDir::new().unwrap();
let mut one_by_one = make_test_volume(tmp2.path().to_str().unwrap());
for id in 1..=10u64 {
let mut n = batch_needle(id, 0xaa, format!("body-{id}").as_bytes());
one_by_one.write_needle(&mut n, true, true).unwrap();
}
assert_eq!(one_by_one.sync_counts_for_test(), (10, 10));
}
/// A failed shared sync fails every entry of the run, durable or not,
/// and leaves the volume as it was before the run: the .dat back at the
/// run start, the clocks where they were, nothing published.
#[test]
fn test_grouped_failed_sync_rolls_back_the_run() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut kept = batch_needle(1, 0xaa, b"first-copy");
v.write_needle(&mut kept, true, true).unwrap();
let prior = v.nm.as_ref().unwrap().get(NeedleId(1)).unwrap().unwrap();
let dat_len_before = dat_len(&v);
let last_append_before = v.last_append_at_ns;
let last_modified_before = v.last_modified_ts_seconds;
let file_count_before = v.file_count();
let mut writes = vec![
(batch_needle(1, 0xaa, b"second-copy"), false),
(batch_needle(2, 0xbb, b"durable"), true),
(batch_needle(3, 0xcc, b"lazy"), false),
];
for (n, _) in writes.iter_mut() {
n.last_modified = last_modified_before + 1000;
}
v.fail_next_fsync_for_test(true);
let results = v.write_needles_grouped(&mut writes);
v.fail_next_fsync_for_test(false);
assert!(
results.iter().all(|r| matches!(r, Err(VolumeError::Io(_)))),
"every entry of the run shares its failed sync: {results:?}"
);
assert_eq!(dat_len(&v), dat_len_before, "the run is off the .dat");
assert_eq!(v.last_append_at_ns, last_append_before);
assert_eq!(v.last_modified_ts_seconds, last_modified_before);
let now = v.nm.as_ref().unwrap().get(NeedleId(1)).unwrap().unwrap();
assert_eq!((now.offset, now.size), (prior.offset, prior.size));
assert!(v.nm.as_ref().unwrap().get(NeedleId(2)).unwrap().is_none());
assert!(v.nm.as_ref().unwrap().get(NeedleId(3)).unwrap().is_none());
assert_eq!(v.file_count(), file_count_before);
assert!(
!v.is_read_only(),
"a rolled-back run keeps the volume writable"
);
let mut read_n = Needle {
id: NeedleId(1),
..Needle::default()
};
v.read_needle(&mut read_n).unwrap();
assert_eq!(read_n.data, b"first-copy");
// The same run goes through once the disk recovers.
let results = v.write_needles_grouped(&mut writes);
assert!(results.iter().all(|r| r.is_ok()));
assert!(v.last_append_at_ns > last_append_before);
}
/// A repeated id starts a new run, so the second write sees the first
/// one published: a different cookie is refused and the same content
/// dedups, exactly as two sequential writes would.
#[test]
fn test_grouped_repeated_id_behaves_as_sequential_writes() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut writes = vec![
(batch_needle(1, 0xaa, b"one"), true),
(batch_needle(1, 0xbb, b"imposter"), true),
(batch_needle(2, 0xcc, b"two"), true),
(batch_needle(2, 0xcc, b"two"), true),
];
let results = v.write_needles_grouped(&mut writes);
assert!(matches!(results[0], Ok((_, _, false))));
assert!(matches!(results[1], Err(VolumeError::CookieMismatch(0xbb))));
assert!(matches!(results[2], Ok((_, _, false))));
assert!(
matches!(results[3], Ok((0, _, true))),
"the same content dedups against the write just before it"
);
// Runs [1], [1, 2], [2]: one pair of syncs each.
assert_eq!(v.sync_counts_for_test(), (3, 3));
let mut read_n = Needle {
id: NeedleId(1),
..Needle::default()
};
v.read_needle(&mut read_n).unwrap();
assert_eq!(read_n.data, b"one");
}
/// A failed shared .idx sync fails every published entry and stops the
/// volume taking writes, as it does for a single durable write.
#[test]
fn test_grouped_failed_idx_sync_quarantines_volume() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut writes = vec![
(batch_needle(1, 0xaa, b"durable"), true),
(batch_needle(2, 0xbb, b"lazy"), false),
];
v.fail_next_idx_sync_for_test(true);
let results = v.write_needles_grouped(&mut writes);
v.fail_next_idx_sync_for_test(false);
assert!(results.iter().all(|r| matches!(r, Err(VolumeError::Io(_)))));
assert_eq!(v.sync_counts_for_test(), (1, 1));
assert!(v.is_read_only());
}
/// A failed shared sync whose truncate also fails leaves an unverified
/// tail, so the volume fails closed exactly as a single write does.
#[test]
fn test_grouped_failed_rollback_marks_volume_unavailable() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut writes = vec![
(batch_needle(1, 0xaa, b"never"), true),
(batch_needle(2, 0xbb, b"landed"), false),
];
v.fail_next_fsync_for_test(true);
v.fail_next_truncate_for_test(true);
let results = v.write_needles_grouped(&mut writes);
v.fail_next_fsync_for_test(false);
v.fail_next_truncate_for_test(false);
assert!(results.iter().all(|r| r.is_err()));
assert!(v.unavailable_error().is_some());
assert!(v.is_read_only());
assert!(v.should_quarantine());
assert!(v.nm.as_ref().unwrap().get(NeedleId(1)).unwrap().is_none());
let mut later = vec![(batch_needle(3, 0xcc, b"refused"), true)];
let results = v.write_needles_grouped(&mut later);
assert!(matches!(results[0], Err(VolumeError::Unavailable(_))));
}
/// An .idx writer that tears its `tear_at`-th row: half of the row
/// reaches the file, then the write fails.
struct TornIdxWriter {
file: File,
writes: usize,
tear_at: usize,
fail_truncates: bool,
}
impl Write for TornIdxWriter {
fn write(&mut self, buf: &[u8]) -> io::Result<usize> {
self.writes += 1;
if self.writes == self.tear_at {
self.file.write_all(&buf[..buf.len() / 2])?;
return Err(io::Error::other("injected torn .idx write"));
}
self.file.write(buf)
}
fn flush(&mut self) -> io::Result<()> {
self.file.flush()
}
}
impl Seek for TornIdxWriter {
fn seek(&mut self, pos: SeekFrom) -> io::Result<u64> {
self.file.seek(pos)
}
}
impl crate::storage::needle_map::IdxFileWriter for TornIdxWriter {
fn sync_all(&self) -> io::Result<()> {
self.file.sync_all()
}
fn truncate_to(&mut self, len: u64) -> io::Result<()> {
if self.fail_truncates {
return Err(io::Error::other("injected trim failure"));
}
self.file.set_len(len)
}
}
/// A durable entry whose index update fails stops the volume taking
/// writes, as it does when sent on its own, so the entries after it in
/// the run are refused instead of indexed behind a row that may be
/// torn. Entries before it stay acked.
#[test]
fn test_grouped_failed_durable_index_refuses_rest_of_run() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut kept = batch_needle(4, 0xdd, b"kept");
v.write_needle(&mut kept, true, true).unwrap();
let mut other = batch_needle(5, 0xee, b"other");
v.write_needle(&mut other, true, true).unwrap();
let nm = v.nm.as_mut().unwrap();
let idx_len = nm.index_file_size();
let file = OpenOptions::new()
.write(true)
.open(format!("{dir}/1.idx"))
.unwrap();
nm.set_idx_file(
Box::new(TornIdxWriter {
file,
writes: 0,
tear_at: 2,
fail_truncates: false,
}),
idx_len,
);
let mut writes = vec![
(batch_needle(1, 0xaa, b"before"), true),
(batch_needle(2, 0xbb, b"torn"), true),
(batch_needle(3, 0xcc, b"after"), true),
(batch_needle(4, 0xdd, b"kept"), true),
(batch_needle(5, 0xef, b"wrong cookie"), true),
];
let results = v.write_needles_grouped(&mut writes);
// Two setup writes, then the run's one sync of each file.
assert_eq!(v.sync_counts_for_test(), (3, 3));
assert!(v.is_read_only());
drop(v);
let reopened = match Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
) {
Ok(v) => v,
Err(e) => panic!("the volume does not reload: {e}"),
};
for ((n, _), r) in writes.iter().zip(&results) {
if r.is_ok() {
let mut got = Needle {
id: n.id,
..Needle::default()
};
let read = reopened.read_needle(&mut got);
assert!(
read.is_ok() && got.data == n.data,
"acked write {} lost on reload: {read:?}",
n.id.0
);
}
}
assert!(matches!(results[0], Ok((_, _, false))), "{results:?}");
assert!(matches!(results[1], Err(VolumeError::Io(_))), "{results:?}");
for r in &results[2..] {
assert!(
matches!(r, Err(VolumeError::ReadOnly(VolumeId(1)))),
"refused as it would be sent on its own: {results:?}"
);
}
}
/// A torn .idx row is trimmed back, so the row a later write appends
/// still lands aligned and survives a reload.
#[test]
fn test_failed_index_write_is_trimmed() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let nm = v.nm.as_mut().unwrap();
let idx_len = nm.index_file_size();
let file = OpenOptions::new()
.write(true)
.open(format!("{dir}/1.idx"))
.unwrap();
nm.set_idx_file(
Box::new(TornIdxWriter {
file,
writes: 0,
tear_at: 1,
fail_truncates: false,
}),
idx_len,
);
let mut writes = vec![
(batch_needle(1, 0xaa, b"torn"), false),
(batch_needle(2, 0xbb, b"after"), true),
];
let results = v.write_needles_grouped(&mut writes);
assert!(matches!(results[0], Err(VolumeError::Io(_))), "{results:?}");
assert!(matches!(results[1], Ok((_, _, false))), "{results:?}");
drop(v);
let reopened = match Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
) {
Ok(v) => v,
Err(e) => panic!("the volume does not reload: {e}"),
};
let mut got = Needle {
id: NeedleId(2),
..Needle::default()
};
reopened.read_needle(&mut got).unwrap();
assert_eq!(got.data, b"after");
}
/// When the trim of a torn .idx row itself fails, no later row is
/// appended after the torn bytes.
#[test]
fn test_untrimmed_torn_row_refuses_later_appends() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let nm = v.nm.as_mut().unwrap();
let idx_len = nm.index_file_size();
let file = OpenOptions::new()
.write(true)
.open(format!("{dir}/1.idx"))
.unwrap();
nm.set_idx_file(
Box::new(TornIdxWriter {
file,
writes: 0,
tear_at: 1,
fail_truncates: true,
}),
idx_len,
);
let mut writes = vec![
(batch_needle(1, 0xaa, b"torn"), false),
(batch_needle(2, 0xbb, b"after"), false),
];
let results = v.write_needles_grouped(&mut writes);
assert!(matches!(results[0], Err(VolumeError::Io(_))), "{results:?}");
assert!(results[1].is_err(), "{results:?}");
// The second row was never appended behind the torn bytes.
let size = std::fs::metadata(format!("{dir}/1.idx")).unwrap().len();
assert_eq!(size, idx_len + 8);
}
/// The I/O error streak after writing four durable needles, the ones in
/// `failing` with a media error on their append, first one at a time and
/// then as one grouped run. Also returns which grouped writes landed.
#[cfg(any(unix, windows))]
fn io_error_streaks(failing: &[u64]) -> (i32, i32, Vec<bool>) {
let writes = || -> Vec<_> {
(1..=4u64)
.map(|id| {
(
batch_needle(id, 0xaa, format!("body-{id}").as_bytes()),
true,
)
})
.collect()
};
let failing: Vec<_> = failing.iter().map(|&id| NeedleId(id)).collect();
let tmp = TempDir::new().unwrap();
let mut one_by_one = make_test_volume(tmp.path().to_str().unwrap());
one_by_one.fail_append_for_test(&failing);
for (mut n, fsync) in writes() {
let _ = one_by_one.write_needle(&mut n, true, fsync);
}
let tmp2 = TempDir::new().unwrap();
let mut v = make_test_volume(tmp2.path().to_str().unwrap());
v.fail_append_for_test(&failing);
let mut grouped = writes();
let results = v.write_needles_grouped(&mut grouped);
assert!(!v.is_read_only(), "the failed appends were truncated back");
let landed = results
.iter()
.map(|r| matches!(r, Ok((_, _, false))))
.collect();
(
one_by_one.get_io_error_state().1,
v.get_io_error_state().1,
landed,
)
}
/// A run counts I/O errors as the same writes sent one at a time would:
/// a success early in the run must not wipe out the streak that the
/// failed appends after it built up, or the volume escapes quarantine.
#[cfg(any(unix, windows))]
#[test]
fn test_grouped_run_keeps_the_io_error_streak_of_later_appends() {
use crate::storage::io_error::IO_ERROR_TOLERANCE;
let (one_by_one, grouped, landed) = io_error_streaks(&[2, 3, 4]);
assert_eq!(landed, [true, false, false, false]);
assert_eq!(one_by_one, IO_ERROR_TOLERANCE);
assert_eq!(grouped, IO_ERROR_TOLERANCE);
}
/// The other half: a write that lands clears the errors of the appends
/// queued before it, even when another append after it fails, or the
/// run reaches a quarantine that the same writes one at a time do not.
#[cfg(any(unix, windows))]
#[test]
fn test_grouped_run_clears_the_io_error_streak_of_earlier_appends() {
let (one_by_one, grouped, landed) = io_error_streaks(&[1, 2, 4]);
assert_eq!(landed, [false, false, true, false]);
assert_eq!(one_by_one, 1);
assert_eq!(grouped, 1);
}
/// An append whose partial bytes cannot be truncated back leaves the .dat
/// tail unverified, so the volume fails closed rather than let a later
/// append bury them mid-file.
#[test]
fn test_failed_append_rollback_marks_volume_unavailable() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut first = batch_needle(1, 0xaa, b"landed");
v.write_needle(&mut first, true, true).unwrap();
// A read-only .dat handle fails the append and the truncate-back alike.
v.dat_file = Some(File::open(v.file_name(".dat")).unwrap());
let mut n = batch_needle(2, 0xbb, b"never-lands");
v.write_needle(&mut n, true, false).unwrap_err();
assert!(v.unavailable_error().is_some());
assert!(v.is_read_only());
assert!(v.should_quarantine());
assert!(v.nm.as_ref().unwrap().get(NeedleId(2)).unwrap().is_none());
let mut later = batch_needle(3, 0xcc, b"refused");
assert!(matches!(
v.write_needle(&mut later, true, true),
Err(VolumeError::Unavailable(_))
));
}
#[test]
fn test_volume_write_dedup() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"same data".to_vec(),
data_size: 9,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
// Write same needle again — should be unchanged
let mut n2 = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"same data".to_vec(),
data_size: 9,
..Needle::default()
};
n2.checksum = CRC::new(&n2.data);
let (_, _, unchanged) = v.write_needle(&mut n2, true, false).unwrap();
assert!(unchanged);
}
#[test]
fn test_volume_delete() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xbb),
data: b"delete me".to_vec(),
data_size: 9,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
assert_eq!(v.file_count(), 1);
let deleted_size = v
.delete_needle(&mut Needle {
id: NeedleId(1),
cookie: Cookie(0xbb),
..Needle::default()
})
.unwrap();
assert!(deleted_size.0 > 0);
// Additive-only: file_count stays at 1 after delete
assert_eq!(v.file_count(), 1);
assert_eq!(v.deleted_count(), 1);
// Read should fail with Deleted
let mut read_n = Needle {
id: NeedleId(1),
..Needle::default()
};
let err = v.read_needle(&mut read_n).unwrap_err();
assert!(matches!(err, VolumeError::Deleted));
}
// Guard the Rust integrity-check tombstone path against the Go regression
// where verifyDeletedNeedleIntegrity forwarded TombstoneFileSize into the
// needle-size check, mismatched against the on-disk Size=0 header, and
// sent every volume with a trailing deletion read-only on load. The Rust
// check guards its size comparison with !size.is_deleted(); this test
// keeps that guarantee from silently regressing.
#[test]
fn test_check_volume_data_integrity_with_deletion_tombstone() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = make_test_volume(dir);
for i in 1..=3 {
let data = format!("data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.delete_needle(&mut Needle {
id: NeedleId(2),
cookie: Cookie(2),
..Needle::default()
})
.unwrap();
v.sync_to_disk().unwrap();
}
let v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert!(
!v.is_no_write_or_delete(),
"volume should not be read-only after reload with trailing deletion tombstone"
);
}
// Reproduce the delete-traffic TTL bug: deletes append a tombstone to the
// .dat, which moves the file's mtime, and the loader read the TTL clock
// back from that mtime. A volume taking delete traffic therefore had
// is_expired() re-armed for another full TTL on every restart and was
// never reclaimed.
#[test]
fn test_ttl_clock_survives_deletes() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let ttl = crate::storage::needle::ttl::TTL::read("5m").unwrap();
let last_write_ns = (SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap()
.as_secs()
- 2 * 60 * 60)
* 1_000_000_000;
let content_size = {
let mut v = make_ttl_volume(dir, ttl);
let mut written = Vec::new();
for i in 1..=3u64 {
let data = format!("data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
let (offset, _, _) = v.write_needle(&mut n, true, false).unwrap();
written.push((offset, n.size));
}
// More than one tombstone: the scan has to walk back over the whole
// run of them to reach a write.
for i in [2u64, 3] {
v.delete_needle(&mut Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
..Needle::default()
})
.unwrap();
}
v.sync_to_disk().unwrap();
// Backdate the writes on disk so the last one sits well outside the
// TTL, while the tombstones leave the .dat mtime at now.
for (offset, size) in written {
backdate_append_at_ns(&v.dat_path(), offset, size, last_write_ns);
}
v.content_size()
};
let v = make_ttl_volume(dir, ttl);
assert_eq!(v.last_modified_ts(), last_write_ns / 1_000_000_000);
assert!(
v.is_expired(content_size, 1024 * 1024),
"a TTL volume whose last write is 2h old must be expired after a reload"
);
}
// The one layout where a .dat's order does not track its write order:
// vacuum rewrites it by key, and an overwrite keeps its original, lower key.
// Reading the position rather than the timestamps would recover the
// highest-key needle's older write time and expire the volume before the
// overwrite has lived out its TTL.
#[test]
fn test_ttl_clock_after_vacuum_takes_newest_write() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let ttl = crate::storage::needle::ttl::TTL::read("5m").unwrap();
let now = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap()
.as_secs();
let old_write_ns = (now - 2 * 60 * 60) * 1_000_000_000;
let new_write_ns = (now - 60) * 1_000_000_000;
let content_size = {
let mut v = make_ttl_volume(dir, ttl);
// Needle 1 is overwritten last but sorts first, so vacuum leaves it
// at the head of the .dat with the newest timestamp of the three.
for (id, append_at_ns) in [(2u64, old_write_ns), (3, old_write_ns), (1, new_write_ns)] {
let data = format!("data {}", id);
let mut n = Needle {
id: NeedleId(id),
cookie: Cookie(id as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
let (offset, _, _) = v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
backdate_append_at_ns(&v.dat_path(), offset, n.size, append_at_ns);
}
v.compact_by_index(0, 0, |_| true).unwrap();
v.commit_compact().unwrap();
v.content_size()
};
let v = make_ttl_volume(dir, ttl);
assert_eq!(v.last_modified_ts(), new_write_ns / 1_000_000_000);
assert!(
!v.is_expired(content_size, 1024 * 1024),
"a volume overwritten a minute ago must not be expired after a vacuum and reload"
);
}
// Guard the destroy time an EC volume is reclaimed on: it was recomputed as
// now+TTL every time the .vif was written, so a read-only mark, a tier
// upload or an EC encode handed an already expiring volume another full TTL.
#[test]
fn test_expire_at_sec_counts_from_last_write() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let ttl = crate::storage::needle::ttl::TTL::read("5m").unwrap();
let now = SystemTime::now()
.duration_since(UNIX_EPOCH)
.unwrap()
.as_secs();
let mut v = make_ttl_volume(dir, ttl);
// A volume with nothing written yet has no last write to count from, and
// must not land in 1970 with its data due for destruction on sight.
assert!(v.expire_at_sec() >= now);
v.set_last_modified_ts_for_test(now - 3600);
let want = now - 3600 + ttl.to_seconds();
for pass in 0..2 {
v.save_volume_info().unwrap();
assert_eq!(
v.volume_info.expire_at_sec, want,
".vif save {} moved the destroy time off the last write",
pass
);
}
}
fn make_ttl_volume(dir: &str, ttl: crate::storage::needle::ttl::TTL) -> Volume {
Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec {
ttl: Some(ttl),
..Default::default()
},
)
.unwrap()
}
fn backdate_append_at_ns(dat_path: &str, offset: u64, size: Size, append_at_ns: u64) {
let ts_offset =
offset + NEEDLE_HEADER_SIZE as u64 + size.0 as u64 + NEEDLE_CHECKSUM_SIZE as u64;
let mut f = OpenOptions::new().write(true).open(dat_path).unwrap();
f.seek(SeekFrom::Start(ts_offset)).unwrap();
f.write_all(&append_at_ns.to_be_bytes()).unwrap();
}
fn write_three_needles(dir: &str) {
let mut v = make_test_volume(dir);
for i in 1..=3u64 {
let data = format!("data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
}
fn reload_volume(dir: &str) -> Volume {
Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap()
}
// A torn shutdown can leave bytes in the .dat past the last indexed
// needle. Appending after such a tail would place the next needle at an
// offset the 8-byte-unit .idx encoding may not represent, so the load
// check must quarantine the volume read-only, matching Go's
// verifyNeedleIntegrity "actual N bytes expected M bytes" behavior.
#[test]
fn test_integrity_detects_unindexed_dat_tail() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
write_three_needles(dir);
let dat_path = volume_file_name(dir, "", VolumeId(1)) + ".dat";
let mut f = OpenOptions::new().append(true).open(&dat_path).unwrap();
f.write_all(&[0xFFu8; 100]).unwrap();
f.sync_all().unwrap();
drop(f);
let v = reload_volume(dir);
assert!(
v.is_no_write_or_delete(),
"volume with unindexed .dat tail must load read-only"
);
}
// Go's tail check is v3-only, so it serves a v1/v2 volume with an
// unindexed tail read-write. Checking it here flipped whole legacy disks
// read-only on their first boot under this server.
#[test]
fn test_integrity_skips_dat_tail_check_before_v3() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec {
version: VERSION_2,
..Default::default()
},
)
.unwrap();
for i in 1..=3u64 {
let data = format!("data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
}
let dat_path = volume_file_name(dir, "", VolumeId(1)) + ".dat";
let mut f = OpenOptions::new().append(true).open(&dat_path).unwrap();
f.write_all(&[0xFFu8; 96]).unwrap();
f.sync_all().unwrap();
drop(f);
let v = reload_volume(dir);
assert_eq!(v.version(), VERSION_2, "volume should reload as v2");
assert!(
!v.is_no_write_or_delete(),
"v2 volume with unindexed .dat tail must stay writable, as in Go"
);
}
// Same torn tail, but with a deletion tombstone as the last indexed
// record — the tombstone path must also verify the .dat ends at it.
#[test]
fn test_integrity_detects_tail_after_tombstone() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = make_test_volume(dir);
let data = b"data 1".to_vec();
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: data.clone(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.delete_needle(&mut Needle {
id: NeedleId(1),
cookie: Cookie(1),
..Needle::default()
})
.unwrap();
v.sync_to_disk().unwrap();
}
let dat_path = volume_file_name(dir, "", VolumeId(1)) + ".dat";
let mut f = OpenOptions::new().append(true).open(&dat_path).unwrap();
f.write_all(&[0xFFu8; 64]).unwrap();
f.sync_all().unwrap();
drop(f);
let v = reload_volume(dir);
assert!(
v.is_no_write_or_delete(),
"volume with unindexed tail after tombstone must load read-only"
);
}
// A key-sorted .idx (legacy weed fix output) puts the highest-key needle
// last by file position, not the .dat-tail needle. The check must find
// the true tail entry via the max-offset scan instead of falsely flipping
// a healthy volume read-only.
#[test]
fn test_integrity_key_sorted_idx_no_false_readonly() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
write_three_needles(dir);
// Swap the last two .idx entries so the file-position-last entry
// references a mid-file needle.
let idx_path = volume_file_name(dir, "", VolumeId(1)) + ".idx";
let content = fs::read(&idx_path).unwrap();
assert_eq!(content.len(), 3 * NEEDLE_MAP_ENTRY_SIZE);
let mut swapped = content.clone();
swapped[NEEDLE_MAP_ENTRY_SIZE..2 * NEEDLE_MAP_ENTRY_SIZE]
.copy_from_slice(&content[2 * NEEDLE_MAP_ENTRY_SIZE..]);
swapped[2 * NEEDLE_MAP_ENTRY_SIZE..]
.copy_from_slice(&content[NEEDLE_MAP_ENTRY_SIZE..2 * NEEDLE_MAP_ENTRY_SIZE]);
fs::write(&idx_path, swapped).unwrap();
let v = reload_volume(dir);
assert!(
!v.is_no_write_or_delete(),
"healthy volume with key-sorted .idx must not load read-only"
);
}
// A .dat truncated mid-needle (crash during append) must quarantine.
#[test]
fn test_integrity_detects_truncated_dat_tail() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
write_three_needles(dir);
let dat_path = volume_file_name(dir, "", VolumeId(1)) + ".dat";
let len = fs::metadata(&dat_path).unwrap().len();
let f = OpenOptions::new().write(true).open(&dat_path).unwrap();
f.set_len(len - 5).unwrap();
f.sync_all().unwrap();
drop(f);
let v = reload_volume(dir);
assert!(
v.is_no_write_or_delete(),
"volume with truncated .dat tail must load read-only"
);
}
#[test]
fn test_scrub_empty_volume() {
// Mirror of Go's TestScrubVolumeData "zero-size volume without index"
// case (weed/storage/volume_checking_test.go): a freshly created /
// pre-allocated volume has a superblock-only .dat and a zero-size .idx,
// and must scrub clean instead of being flagged as corrupt.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let v = make_test_volume(dir);
// .dat holds only the superblock; .idx is empty.
assert_eq!(v.dat_file_size().unwrap(), SUPER_BLOCK_SIZE as u64);
let (files_checked, broken) = v.scrub().unwrap();
assert_eq!(files_checked, 0);
assert!(
broken.is_empty(),
"empty volume should scrub clean, got {:?}",
broken
);
// The index-only mode must agree.
let (idx_checked, idx_broken) = v.scrub_index().unwrap();
assert_eq!(idx_checked, 0);
assert!(
idx_broken.is_empty(),
"empty volume should scrub_index clean, got {:?}",
idx_broken
);
}
#[test]
fn test_scrub_overwritten_volume_scrubs_clean() {
// An overwrite/delete appends a second .dat record + .idx row. Walking
// the on-disk .idx (not the deduped map) makes total_read match the
// physical .dat, so a healthy overwritten volume is NOT flagged broken.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for (id, body) in [(1u64, "first version"), (1, "second version"), (2, "other")] {
let mut n = Needle {
id: NeedleId(id),
cookie: Cookie(id as u32),
data: body.as_bytes().to_vec(),
data_size: body.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
let (count, broken) = v.scrub().unwrap();
assert!(
broken.is_empty(),
"overwritten volume should scrub clean, got {:?}",
broken
);
assert_eq!(count, 3, "scrub must count every .idx row");
}
#[test]
fn test_scrub_ignores_offset0_tombstone() {
// A remote-tier delete appends an offset-0 tombstone to the .idx with no
// physical .dat bytes; it must not flag the volume (neither the size
// reconcile nor the structural overlap check).
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let data = b"needle data".to_vec();
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: data.clone(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
// Append an offset-0 logical tombstone to the on-disk .idx.
let mut f = std::fs::OpenOptions::new()
.append(true)
.open(v.file_name(".idx"))
.unwrap();
crate::storage::idx::write_index_entry(
&mut f,
NeedleId(2),
Offset::from_actual_offset(0),
Size(-1),
)
.unwrap();
f.sync_all().unwrap();
drop(f);
let (count, broken) = v.scrub().unwrap();
assert!(
broken.is_empty(),
"offset-0 tombstone must not flag the volume, got {:?}",
broken
);
assert_eq!(count, 2, "both .idx rows are walked");
}
/// The .dat offset a local delete's .idx row points at: the physical
/// tombstone record. Mirrors the Go test helper localTombstoneOffset.
fn local_tombstone_offset(v: &Volume, id: u64) -> Offset {
let mut idx_file = File::open(v.file_name(".idx")).unwrap();
let mut found = Offset::default();
idx::walk_index_file(&mut idx_file, 0, |key, offset, size| {
if key == NeedleId(id) && !offset.is_zero() && size.is_deleted() {
found = offset;
}
Ok(())
})
.unwrap();
assert!(
!found.is_zero(),
"expected a local deletion tombstone for needle {} in .idx",
id
);
found
}
#[test]
fn test_scrub_checks_local_deletion_tombstone() {
// Mirror of Go's TestScrubVolumeDataChecksLocalDeletionTombstone: a
// local delete's tombstone record is read and its needle id checked
// against the .idx key.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, b"needle data");
v.delete_needle(&mut Needle {
id: NeedleId(1),
cookie: Cookie(0x12345678),
..Needle::default()
})
.unwrap();
// A later record keeps the tombstone mid-file rather than at the tail.
write_test_needle(&mut v, 2, b"needle data");
v.sync_to_disk().unwrap();
let tombstone_offset = local_tombstone_offset(&v, 1);
let (_count, broken) = v.scrub().unwrap();
assert!(
broken.is_empty(),
"healthy local deletion tombstone must pass scrub, got {:?}",
broken
);
let mut id_bytes = [0u8; NEEDLE_ID_SIZE];
NeedleId(99).to_bytes(&mut id_bytes);
let mut dat = OpenOptions::new()
.write(true)
.open(v.file_name(".dat"))
.unwrap();
dat.seek(SeekFrom::Start(
tombstone_offset.to_actual_offset() as u64 + COOKIE_SIZE as u64,
))
.unwrap();
dat.write_all(&id_bytes).unwrap();
dat.sync_all().unwrap();
let (_count, broken) = v.scrub().unwrap();
assert!(
broken
.iter()
.any(|e| e.contains("does not match needle's Id")),
"scrub should report the corrupted tombstone's Id, got {:?}",
broken
);
}
#[test]
fn test_scrub_checks_live_needle_id() {
// Mirror of Go's TestScrubVolumeDataChecksLiveNeedleId: a live needle
// whose stored id is damaged still reads clean — the data CRC does not
// cover the header — so the scrub must catch it via the index key.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, b"needle data");
v.sync_to_disk().unwrap();
let mut idx_file = File::open(v.file_name(".idx")).unwrap();
let mut live_offset = Offset::default();
idx::walk_index_file(&mut idx_file, 0, |key, offset, size| {
if key == NeedleId(1) && !offset.is_zero() && !size.is_deleted() {
live_offset = offset;
}
Ok(())
})
.unwrap();
let (_count, broken) = v.scrub().unwrap();
assert!(
broken.is_empty(),
"healthy needle must pass, got {:?}",
broken
);
let mut id_bytes = [0u8; NEEDLE_ID_SIZE];
NeedleId(99).to_bytes(&mut id_bytes);
let mut dat = OpenOptions::new()
.write(true)
.open(v.file_name(".dat"))
.unwrap();
dat.seek(SeekFrom::Start(
live_offset.to_actual_offset() as u64 + COOKIE_SIZE as u64,
))
.unwrap();
dat.write_all(&id_bytes).unwrap();
dat.sync_all().unwrap();
let (_count, broken) = v.scrub().unwrap();
assert!(
broken
.iter()
.any(|e| e.contains("does not match needle's Id")),
"scrub should report the corrupted live needle's Id, got {:?}",
broken
);
}
#[test]
fn test_scrub_reports_truncated_local_deletion_tombstone() {
// Mirror of Go's TestScrubVolumeDataReportsTruncatedLocalDeletionTombstone.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, b"needle data");
v.delete_needle(&mut Needle {
id: NeedleId(1),
cookie: Cookie(0x12345678),
..Needle::default()
})
.unwrap();
v.sync_to_disk().unwrap();
let tombstone_offset = local_tombstone_offset(&v, 1);
let dat = OpenOptions::new()
.write(true)
.open(v.file_name(".dat"))
.unwrap();
dat.set_len(tombstone_offset.to_actual_offset() as u64 + NEEDLE_HEADER_SIZE as u64)
.unwrap();
dat.sync_all().unwrap();
let (_count, broken) = v.scrub().unwrap();
assert!(
broken
.iter()
.any(|e| e.contains("failed to read needle 1 on volume 1")),
"scrub should report the truncated tombstone read, got {:?}",
broken
);
}
#[test]
fn test_scrub_index_flags_zero_size_idx_with_data() {
// A populated .dat with an empty .idx is corruption — the pre-allocated
// exception only covers a superblock-only .dat.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let data = b"needle data".to_vec();
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: data.clone(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
assert!(v.dat_file_size().unwrap() > SUPER_BLOCK_SIZE as u64);
// Truncate the .idx to zero while the .dat keeps its needle.
std::fs::File::create(v.file_name(".idx")).unwrap();
let (count, broken) = v.scrub_index().unwrap();
assert_eq!(count, 0);
assert!(
broken.iter().any(|e| e.contains("zero-size IDX file")),
"expected zero-size IDX error, got {:?}",
broken
);
}
#[test]
fn test_scrub_healthy_volume() {
// Mirror of Go's TestScrubVolumeData "healthy volume" case: a volume
// with live needles scrubs clean and the .dat size accounting matches.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for i in 1..=5 {
let data = format!("needle data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
let (files_checked, broken) = v.scrub().unwrap();
assert_eq!(files_checked, 5);
assert!(
broken.is_empty(),
"healthy volume should scrub clean, got {:?}",
broken
);
}
#[test]
fn test_volume_multiple_needles() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for i in 1..=10 {
let data = format!("needle data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
assert_eq!(v.file_count(), 10);
assert_eq!(v.max_file_key(), NeedleId(10));
// Read back needle 5
let mut n = Needle {
id: NeedleId(5),
..Needle::default()
};
v.read_needle(&mut n).unwrap();
assert_eq!(n.data, b"needle data 5");
}
#[test]
fn test_volume_reload_from_disk() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
// Write some needles
{
let mut v = make_test_volume(dir);
for i in 1..=3 {
let data = format!("data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
}
// Reload and verify
let v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert_eq!(v.file_count(), 3);
let mut n = Needle {
id: NeedleId(2),
..Needle::default()
};
v.read_needle(&mut n).unwrap();
assert_eq!(std::str::from_utf8(&n.data).unwrap(), "data 2");
}
#[test]
fn test_redb_volume_close_then_reload_keeps_counters_exact() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let open = || {
Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::Redb,
&VolumeSpec::default(),
)
.unwrap()
};
{
let mut v = open();
for i in 1..=3 {
let data = format!("data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
// Volume::close is the shutdown path (store -> disk location ->
// volume). It must checkpoint the redb index, or the reload
// replays rows the table already holds and inflates the counters.
v.close();
}
let v = open();
assert_eq!(v.file_count(), 3);
assert_eq!(v.deleted_count(), 0);
let mut n = Needle {
id: NeedleId(2),
..Needle::default()
};
v.read_needle(&mut n).unwrap();
assert_eq!(std::str::from_utf8(&n.data).unwrap(), "data 2");
}
#[test]
fn test_redb_volume_checkpoint_flushes_dat_before_index() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::Redb,
&VolumeSpec::default(),
)
.unwrap();
let write = |v: &mut Volume, i: u64| {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(1),
data: b"x".to_vec(),
data_size: 1,
..Needle::default()
};
// fsync=false: the .dat append is not flushed by the write itself.
v.write_needle(&mut n, true, false).unwrap();
};
for i in 1..1000 {
write(&mut v, i);
}
// No checkpoint due yet, so META still holds the load-time .idx
// size. Read live: copying the open .rdb fails on Windows, where
// redb's file lock is mandatory (see live_meta_idx_size).
assert_eq!(v.live_meta_idx_size_for_test(), Some(0));
// The 1000th write makes an index checkpoint due. A checkpoint makes
// the index durable, so the volume has to flush the .dat first, and
// when that flush fails nothing may be checkpointed: a durable index
// row pointing past the end of an unflushed .dat loads read-only.
v.fail_next_fsync_for_test(true);
write(&mut v, 1000);
assert_eq!(
v.live_meta_idx_size_for_test(),
Some(0),
"skipped checkpoint must not record progress"
);
v.fail_next_fsync_for_test(false);
write(&mut v, 1001);
assert_eq!(
v.live_meta_idx_size_for_test(),
Some(1001 * NEEDLE_MAP_ENTRY_SIZE as u64)
);
}
#[test]
fn test_relocate_index_to_moves_index_and_serves_reads() {
let root = TempDir::new().unwrap();
let data_dir = root.path().join("data");
let idx_dir = root.path().join("idx");
fs::create_dir_all(&data_dir).unwrap();
fs::create_dir_all(&idx_dir).unwrap();
let data = data_dir.to_str().unwrap();
let idx = idx_dir.to_str().unwrap();
// Create the volume with its index co-located in the data dir (the state
// an EC reconstruct leaves), write a needle, and flush.
let mut v = Volume::new(
data,
data,
VolumeId(7),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
let payload = b"payload-across-relocate".to_vec();
let mut n = Needle {
id: NeedleId(42),
cookie: Cookie(0x55),
data: payload.clone(),
data_size: payload.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
let data_idx = format!("{data}/7.idx");
let idx_dir_idx = format!("{idx}/7.idx");
assert!(
Path::new(&data_idx).exists(),
"precondition: index co-located with the data"
);
v.relocate_index_to(idx).unwrap();
assert!(
Path::new(&idx_dir_idx).exists(),
"index moved to the idx dir"
);
assert!(
!Path::new(&data_idx).exists(),
"index gone from the data dir"
);
assert_eq!(v.file_count(), 1);
// The volume reloaded in place, so the needle still reads back.
let mut got = Needle {
id: NeedleId(42),
..Needle::default()
};
v.read_needle(&mut got).unwrap();
assert_eq!(got.data, payload, "needle content survives the relocate");
// Relocating again is a no-op: the index is already where asked.
v.relocate_index_to(idx).unwrap();
assert!(Path::new(&idx_dir_idx).exists());
}
#[test]
fn test_load_keeps_index_co_located_with_the_data() {
let root = TempDir::new().unwrap();
let data_dir = root.path().join("data");
let idx_dir = root.path().join("idx");
fs::create_dir_all(&data_dir).unwrap();
fs::create_dir_all(&idx_dir).unwrap();
let data = data_dir.to_str().unwrap();
let idx = idx_dir.to_str().unwrap();
let mut v = Volume::new(
data,
data,
VolumeId(7),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
let payload = b"payload-beside-the-data".to_vec();
let mut n = Needle {
id: NeedleId(42),
cookie: Cookie(0x55),
data: payload.clone(),
data_size: payload.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
drop(v);
// --dir.idx now names an empty directory: the index already beside the
// data keeps serving, and nothing lands in the new directory.
let reopened = Volume::new(
data,
idx,
VolumeId(7),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert!(
Path::new(&format!("{data}/7.idx")).exists(),
"index stays with the data"
);
assert!(
!Path::new(&format!("{idx}/7.idx")).exists(),
"nothing written to the new idx dir"
);
let mut got = Needle {
id: NeedleId(42),
..Needle::default()
};
reopened.read_needle(&mut got).unwrap();
assert_eq!(got.data, payload);
}
#[test]
fn test_relocate_index_to_noop_when_already_in_place() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: b"x".to_vec(),
data_size: 1,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
// Data and index share a directory, so there is nothing to move.
v.relocate_index_to(dir).unwrap();
assert!(Path::new(&format!("{dir}/1.idx")).exists());
let mut got = Needle {
id: NeedleId(1),
..Needle::default()
};
v.read_needle(&mut got).unwrap();
assert_eq!(got.data, b"x");
}
#[test]
fn test_volume_cookie_mismatch() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"original".to_vec(),
data_size: 8,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
// Write with wrong cookie
let mut n2 = Needle {
id: NeedleId(1),
cookie: Cookie(0xbb),
data: b"overwrite".to_vec(),
data_size: 9,
..Needle::default()
};
let err = v.write_needle(&mut n2, true, false).unwrap_err();
assert!(matches!(err, VolumeError::CookieMismatch(_)));
}
// A size disagreeing with the blob's own header indexes the needle at the wrong
// length and, on v3, stamps the append timestamp into the middle of the needle.
#[test]
fn test_write_needle_blob_rejects_size_mismatch() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0x12345678),
data: b"the merged payload".to_vec(),
data_size: 18,
..Needle::default()
};
n.checksum = CRC::new(&n.data);
let (offset, _, _) = v.write_needle(&mut n, true, false).unwrap();
let blob = v.read_needle_blob(offset as i64, n.size).unwrap();
let dat_size_before = v.dat_file_size().unwrap();
// Size(n.data_size) is what append reports, and what a caller following the
// payload-size convention would send.
let err = v
.write_needle_blob_and_index(NeedleId(2), &blob, Size(n.data_size as i32))
.unwrap_err();
assert!(matches!(err, VolumeError::Io(_)), "got {err:?}");
assert_eq!(v.dat_file_size().unwrap(), dat_size_before);
v.write_needle_blob_and_index(NeedleId(2), &blob, n.size)
.unwrap();
}
// A blob shorter or longer than its record size leaves .dat off the record
// grid: later writes index at truncated offsets, or a scan reads the leftover
// bytes as the next record.
#[test]
fn test_write_needle_blob_rejects_length_mismatch() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0x12345678),
data: b"the merged payload".to_vec(),
data_size: 18,
..Needle::default()
};
n.checksum = CRC::new(&n.data);
let (offset, _, _) = v.write_needle(&mut n, true, false).unwrap();
let blob = v.read_needle_blob(offset as i64, n.size).unwrap();
let dat_size_before = v.dat_file_size().unwrap();
let mut too_long = blob.clone();
too_long.push(0);
let too_short = blob[..blob.len() - 1].to_vec();
let mut eight_long = blob.clone();
eight_long.extend_from_slice(&[0u8; 8]);
for mutated in [&too_long, &too_short, &eight_long] {
let err = v
.write_needle_blob_and_index(NeedleId(2), mutated, n.size)
.unwrap_err();
assert!(matches!(err, VolumeError::Io(_)), "got {err:?}");
assert_eq!(v.dat_file_size().unwrap(), dat_size_before);
}
// Later ordinary writes still read back correctly.
let mut next = Needle {
id: NeedleId(3),
cookie: Cookie(2),
data: b"next".to_vec(),
data_size: 4,
..Needle::default()
};
next.checksum = CRC::new(&next.data);
v.write_needle(&mut next, true, false).unwrap();
let mut got = Needle {
id: NeedleId(3),
..Needle::default()
};
v.read_needle(&mut got).unwrap();
assert_eq!(got.data, b"next");
}
#[test]
fn test_read_blob_negative_does_not_panic() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let v = make_test_volume(dir);
let res = v.read_needle_blob(0, Size(-100));
assert!(res.is_err(), "negative size must return Err, not panic");
match res.unwrap_err() {
VolumeError::Io(e) => assert_eq!(e.kind(), std::io::ErrorKind::InvalidData),
e => panic!("expected Io InvalidData, got {e:?}"),
}
}
#[test]
fn test_read_blob_zero_size_not_rejected_by_validation() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let v = make_test_volume(dir);
// Size(0) must not be rejected by the negativity guard: the read may
// Ok or fail on empty-volume IO, but never with our validation message.
match v.read_needle_blob(0, Size(0)) {
Ok(_) => {}
Err(VolumeError::Io(e)) => {
assert!(
!e.to_string().contains("invalid needle size"),
"Size(0) must pass validation, got {e}"
);
}
Err(e) => panic!("unexpected error kind for Size(0): {e:?}"),
}
}
#[test]
fn test_volume_destroy() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let dat_path;
let idx_path;
{
let mut v = make_test_volume(dir);
dat_path = v.file_name(".dat");
idx_path = v.file_name(".idx");
assert!(Path::new(&dat_path).exists());
v.destroy(false, false, false).unwrap();
}
assert!(!Path::new(&dat_path).exists());
assert!(!Path::new(&idx_path).exists());
}
#[test]
fn test_get_append_at_ns() {
let t1 = get_append_at_ns(0);
assert!(t1 > 0);
let t2 = get_append_at_ns(t1);
assert!(t2 > t1);
// If we pass a future timestamp, should return last+1
let future = u64::MAX - 1;
let t3 = get_append_at_ns(future);
assert_eq!(t3, future + 1);
}
#[test]
fn test_volume_compact() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
// Write 3 needles
for i in 1..=3u64 {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: format!("data-{}", i).into_bytes(),
data_size: format!("data-{}", i).len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
assert_eq!(v.file_count(), 3);
// Delete needle 2
let mut del = Needle {
id: NeedleId(2),
cookie: Cookie(2),
..Needle::default()
};
v.delete_needle(&mut del).unwrap();
// Additive-only: file_count stays at 3 after delete
assert_eq!(v.file_count(), 3);
assert_eq!(v.deleted_count(), 1);
let dat_size_before = v.dat_file_size().unwrap();
// Compact
v.compact_by_index(0, 0, |_| true).unwrap();
// Verify compact files exist
assert!(Path::new(&v.file_name(".cpd")).exists());
assert!(Path::new(&v.file_name(".cpx")).exists());
// Commit: swap files and reload
v.commit_compact().unwrap();
// After compaction: 2 live needles, 0 deleted
assert_eq!(v.file_count(), 2);
assert_eq!(v.deleted_count(), 0);
// Dat should be smaller (deleted needle removed)
let dat_size_after = v.dat_file_size().unwrap();
assert!(
dat_size_after < dat_size_before,
"dat should shrink after compact"
);
// Read back live needles
let mut n1 = Needle {
id: NeedleId(1),
..Needle::default()
};
v.read_needle(&mut n1).unwrap();
assert_eq!(n1.data, b"data-1");
let mut n3 = Needle {
id: NeedleId(3),
..Needle::default()
};
v.read_needle(&mut n3).unwrap();
assert_eq!(n3.data, b"data-3");
// Needle 2 should not exist
let mut n2 = Needle {
id: NeedleId(2),
..Needle::default()
};
assert!(v.read_needle(&mut n2).is_err());
// Compact files should not exist after commit
assert!(!Path::new(&v.file_name(".cpd")).exists());
assert!(!Path::new(&v.file_name(".cpx")).exists());
// Cleanup should be a no-op
v.cleanup_compact().unwrap();
}
/// Guards the copy-phase integrity check against regressing into
/// double-subtracting skipped bytes: expected_live_bytes already excludes
/// needles dropped as unreadable (they continue before the tally), so the
/// check must compare it directly against the compacted .dat size.
#[test]
fn test_exceeds_expected_compacted_size() {
assert!(!exceeds_expected_compacted_size(100, 100));
assert!(!exceeds_expected_compacted_size(100, 150));
assert!(exceeds_expected_compacted_size(100, 90));
}
/// A write that lands on the live volume mid-copy must not trip the
/// post-copy integrity check. The Rust port snapshots the index entries
/// before the copy loop (equivalent to Go's frozen oldNm), so a concurrent
/// write is invisible to the tally and the check stays quiet. The write
/// is then replayed by makeup_diff during commit_compact.
#[test]
fn test_compact_by_index_tolerates_concurrent_write() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for i in 1..=8u64 {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: format!("data-{}", i).into_bytes(),
data_size: format!("data-{}", i).len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.compact_by_index(0, 0, |_| true).unwrap();
assert!(Path::new(&v.file_name(".cpd")).exists());
// A write arriving after the snapshot but before commit must survive
// via makeup_diff, exactly like a concurrent write in the Go server.
let mut late = Needle {
id: NeedleId(99),
cookie: Cookie(99),
data: b"late-write".to_vec(),
data_size: 10,
..Needle::default()
};
v.write_needle(&mut late, true, false).unwrap();
v.commit_compact().unwrap();
let mut got = Needle {
id: NeedleId(99),
cookie: Cookie(99),
..Needle::default()
};
v.read_needle(&mut got).unwrap();
assert_eq!(got.data, b"late-write");
}
/// Mirrors Go's TestConcurrentWriteCrossesOffsetBoundary (issue #11410):
/// a replayed write's index entry must encode all offset bytes even when
/// old and new offsets sit in different 32 GiB ranges. Unix-only: Windows
/// set_len eagerly allocates the 64 GiB extension.
#[test]
#[cfg(all(unix, feature = "5bytes"))]
fn test_makeup_diff_replay_crosses_offset_boundary() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
write_test_needle(&mut v, 1, &[b'x'; 32]);
v.dat_file.as_ref().unwrap().set_len(64u64 << 30).unwrap();
v.compact_by_index(0, 0, |_| true).unwrap();
let late_data = vec![b'x'; 113];
write_test_needle(&mut v, 342511246, &late_data);
let mut before = Needle {
id: NeedleId(342511246),
..Needle::default()
};
v.read_needle(&mut before).unwrap();
assert_eq!(before.data, late_data);
let expected_offset = fs::metadata(v.file_name(".cpd")).unwrap().len() as i64;
v.commit_compact().unwrap();
let mut idx_file = File::open(v.file_name(".idx")).unwrap();
let mut actual_offset = 0i64;
crate::storage::idx::walk_index_file(&mut idx_file, 0, |key, offset, _size| {
if key == NeedleId(342511246) {
actual_offset = offset.to_actual_offset();
}
Ok(())
})
.unwrap();
assert_eq!(actual_offset, expected_offset);
let mut after = Needle {
id: NeedleId(342511246),
..Needle::default()
};
v.read_needle(&mut after).unwrap();
assert_eq!(after.data, late_data);
v.compact_by_index(0, 0, |_| true).unwrap();
v.commit_compact().unwrap();
let mut third = Needle {
id: NeedleId(342511246),
..Needle::default()
};
v.read_needle(&mut third).unwrap();
assert_eq!(third.data, late_data);
}
/// Vacuum compaction must tolerate an .idx entry whose offset points past
/// the end of the .dat file. The bad
/// entry is silently dropped from the resulting .cpx; healthy needles
/// survive untouched.
#[test]
fn test_compact_by_index_drops_dangling_needle() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
// Write a handful of healthy needles to establish a baseline .dat/.idx.
for i in 1..=5u64 {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: format!("payload-{}", i).into_bytes(),
data_size: format!("payload-{}", i).len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
let dat_size = v.dat_file_size().unwrap();
let bad_key = NeedleId(9999);
let bad_offset = Offset::from_actual_offset((dat_size + 1024 * 1024) as i64);
let bad_size = Size(2048);
v.nm.as_mut()
.expect("needle map present")
.put(bad_key, bad_offset, bad_size)
.unwrap();
v.nm.as_ref().unwrap().sync().unwrap();
// Vacuum must succeed in spite of the dangling entry.
v.compact_by_index(0, 0, |_| true)
.expect("compact_by_index should tolerate dangling entries");
// Walk the resulting .cpx and confirm the bad key was dropped while
// every healthy key made it through.
let cpx_path = v.file_name(".cpx");
let mut cpx = File::open(&cpx_path).unwrap();
let mut kept: Vec<u64> = Vec::new();
idx::walk_index_file(&mut cpx, 0, |key, _, size| {
if !size.is_deleted() {
kept.push(key.0);
}
Ok(())
})
.unwrap();
assert!(
!kept.contains(&bad_key.0),
"dangling key {} should have been dropped from .cpx, got {:?}",
bad_key.0,
kept
);
for i in 1..=5u64 {
assert!(
kept.contains(&i),
"healthy key {} missing from compacted .cpx, got {:?}",
i,
kept
);
}
}
/// The needle map's max_needle_end accumulator must let volume.load
/// detect an .idx whose entries point past the end of the .dat — the
/// deeper-than-tail corruption shape that the existing
/// last-10-entries scan cannot see. The check is populated by the load
/// walk and read in volume.load() to flip the volume read-only.
#[test]
fn test_max_needle_end_detects_dangling_entry() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for i in 1..=4u64 {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: format!("data-{}", i).into_bytes(),
data_size: format!("data-{}", i).len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
let dat_size = v.dat_file_size().unwrap() as i64;
let idx_path = v.file_name(".idx");
let version = v.version();
// Sanity: a fresh load walk over the healthy .idx puts max_needle_end
// somewhere inside the .dat.
let mut idx_reader = File::open(&idx_path).unwrap();
let healthy_nm = CompactNeedleMap::load_from_idx(&mut idx_reader, version).unwrap();
let healthy_end = healthy_nm.max_needle_end();
assert!(
healthy_end > 0 && healthy_end <= dat_size,
"healthy volume should have max_needle_end ({}) in [1, dat_size={}]",
healthy_end,
dat_size
);
// Inject a dangling entry by appending a bogus 16/17-byte record
// directly to the .idx, then reload. The load walk should observe
// max_needle_end past dat_size — which is exactly the signal
// volume.load uses to mark the volume read-only.
let bad_offset = Offset::from_actual_offset(dat_size + 4 * 1024 * 1024);
let mut idx_append = OpenOptions::new().append(true).open(&idx_path).unwrap();
idx::write_index_entry(&mut idx_append, NeedleId(9999), bad_offset, Size(1024)).unwrap();
idx_append.sync_all().unwrap();
let mut idx_reread = File::open(&idx_path).unwrap();
let bad_nm = CompactNeedleMap::load_from_idx(&mut idx_reread, version).unwrap();
let bad_end = bad_nm.max_needle_end();
assert!(
bad_end > dat_size,
"after dangling-entry inject max_needle_end ({}) should exceed dat_size ({})",
bad_end,
dat_size
);
}
#[test]
fn test_stream_info_locates_needle_data() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
// Write a needle large enough to have meaningful data
let data = vec![0xABu8; 2048];
let mut n = Needle {
id: NeedleId(42),
cookie: Cookie(0xDEADBEEF),
data: data.clone(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
// Read stream info
let mut read_n = Needle {
id: NeedleId(42),
cookie: Cookie(0xDEADBEEF),
..Needle::default()
};
let mut plan = v.needle_read_plan(NeedleId(42), false).unwrap();
plan.read_meta(&mut read_n).unwrap();
let info = plan.into_stream_info(&read_n);
assert_eq!(info.needle_id, NeedleId(42));
assert_eq!(info.data_size, data.len() as u32);
assert!(info.data_file_offset > 0);
}
#[test]
fn test_remote_vif_load_blocks_writes_but_allows_delete() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let dat_size_before_reload = {
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0x1234),
data: b"remote".to_vec(),
data_size: 6,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
let vif = VifVolumeInfo {
files: vec![VifRemoteFile {
backend_type: "s3".to_string(),
backend_id: "vif_rw_test".to_string(),
key: "remote-key".to_string(),
offset: 0,
file_size: v.dat_file_size().unwrap(),
modified_time: 123,
extension: ".dat".to_string(),
}],
version: v.version().0 as u32,
..VifVolumeInfo::default()
};
std::fs::write(
format!("{}/1.vif", dir),
serde_json::to_string_pretty(&vif).unwrap(),
)
.unwrap();
// Register a dummy S3 backend so load_remote_dat_file can look it
// up. This test never reads remote data, so a dead endpoint is fine.
// Use a test-specific backend_id to avoid racing with other tests
// that share the global registry.
let tier_config = crate::remote_storage::s3_tier::S3TierConfig {
access_key: "access".to_string(),
secret_key: "secret".to_string(),
region: "us-east-1".to_string(),
bucket: "bucket-a".to_string(),
endpoint: "http://127.0.0.1:1".to_string(),
storage_class: "STANDARD".to_string(),
force_path_style: true,
};
crate::remote_storage::s3_tier::global_s3_tier_registry()
.write()
.unwrap()
.register(
"s3.vif_rw_test".to_string(),
crate::remote_storage::s3_tier::S3TierBackend::new(&tier_config),
);
v.dat_file_size().unwrap()
};
let mut v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert!(v.is_read_only());
assert!(!v.no_write_or_delete);
assert!(v.no_write_can_delete);
let err = v
.write_needle(
&mut Needle {
id: NeedleId(2),
cookie: Cookie(0x5678),
data: b"blocked".to_vec(),
data_size: 7,
..Needle::default()
},
true,
false,
)
.unwrap_err();
assert!(matches!(err, VolumeError::ReadOnly(_)));
let deleted_size = v
.delete_needle(&mut Needle {
id: NeedleId(1),
cookie: Cookie(0x1234),
..Needle::default()
})
.unwrap();
assert!(deleted_size.0 > 0);
assert_eq!(v.dat_file_size().unwrap(), dat_size_before_reload);
crate::remote_storage::s3_tier::global_s3_tier_registry()
.write()
.unwrap()
.remove("s3.vif_rw_test");
}
/// Build a volume with `needles` needles, tier it to a fake remote backend,
/// and reload it the way a restart would — so it comes back read-only with
/// its index on disk. The local .dat is left in place, matching the state a
/// completed tier-down download leaves behind.
fn reload_as_tiered(dir: &str, backend_id: &str, needles: u64) -> Volume {
{
let mut v = make_test_volume(dir);
for i in 1..=needles {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(0x1234),
data: format!("needle-{i}").into_bytes(),
data_size: format!("needle-{i}").len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
let vif = VifVolumeInfo {
files: vec![VifRemoteFile {
backend_type: "s3".to_string(),
backend_id: backend_id.to_string(),
key: "remote-key".to_string(),
offset: 0,
file_size: v.dat_file_size().unwrap(),
modified_time: 123,
extension: ".dat".to_string(),
}],
version: v.version().0 as u32,
..VifVolumeInfo::default()
};
std::fs::write(
format!("{}/1.vif", dir),
serde_json::to_string_pretty(&vif).unwrap(),
)
.unwrap();
let tier_config = crate::remote_storage::s3_tier::S3TierConfig {
access_key: "access".to_string(),
secret_key: "secret".to_string(),
region: "us-east-1".to_string(),
bucket: "bucket-a".to_string(),
endpoint: "http://127.0.0.1:1".to_string(),
storage_class: "STANDARD".to_string(),
force_path_style: true,
};
crate::remote_storage::s3_tier::global_s3_tier_registry()
.write()
.unwrap()
.register(
format!("s3.{backend_id}"),
crate::remote_storage::s3_tier::S3TierBackend::new(&tier_config),
);
}
Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap()
}
/// The remote reference and the derived write mode are the same fact, so a
/// caller must not be able to move one without the other. Every edit goes
/// through update_remote_files, which refreshes the mode on the way out.
#[test]
fn test_update_remote_files_refreshes_the_derived_write_mode() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let file_size = v.dat_file_size().unwrap();
assert!(!v.has_remote_file());
assert!(!v.is_no_write_can_delete());
assert!(!v.is_no_write_or_delete());
v.update_remote_files(|files| {
files.push(PbRemoteFile {
backend_type: "s3".to_string(),
backend_id: "default".to_string(),
key: "remote-key".to_string(),
offset: 0,
file_size,
modified_time: 123,
extension: ".dat".to_string(),
})
})
.unwrap();
assert!(v.has_remote_file());
assert!(
v.is_no_write_can_delete(),
"a remote-backed volume only serves deletes"
);
assert!(!v.is_no_write_or_delete());
v.update_remote_files(|files| {
files.remove(0);
})
.unwrap();
assert!(!v.has_remote_file());
assert!(
!v.is_no_write_can_delete(),
"dropping the last remote reference publishes the volume as writable"
);
assert!(!v.is_no_write_or_delete());
}
// Tier-down clears the remote mode and publishes the volume as writable. The
// map it booted with is the read-only sorted one, whose put always fails, so
// without a rebuild the first write appends to .dat and then cannot be
// indexed — leaving bytes nothing references.
/// A volume tiered while the server runs must install the sorted map right
/// then. Reconciling only on the way down left it holding the whole index
/// in RAM and its .idx descriptor pinned until the next restart — the cost
/// the sorted map exists to avoid.
#[test]
fn test_tier_up_installs_the_sorted_needle_map() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for i in 1..=2u64 {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(0x1234),
data: format!("needle-{i}").into_bytes(),
data_size: format!("needle-{i}").len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
assert!(
!matches!(v.nm, Some(NeedleMap::SortedFile(_))),
"a local volume starts with a writable in-memory map"
);
// What the tier-up handler does once the .dat is uploaded: record the
// remote reference and reconcile the mode.
let file_size = v.dat_file_size().unwrap();
v.update_remote_files(|files| {
files.push(PbRemoteFile {
backend_type: "s3".to_string(),
backend_id: "vif_tierup_test".to_string(),
key: "remote-key".to_string(),
offset: 0,
file_size,
modified_time: 123,
extension: ".dat".to_string(),
})
})
.unwrap();
assert!(v.has_remote_file());
assert!(
matches!(v.nm, Some(NeedleMap::SortedFile(_))),
"tier-up must install the sorted map without waiting for a restart"
);
// The index still answers, now off .sdx.
let nv =
v.nm.as_ref()
.unwrap()
.get(NeedleId(1))
.unwrap()
.expect("needle 1 still resolves through the sorted index");
assert!(!nv.size.is_deleted());
}
#[test]
fn test_tier_down_swaps_in_a_writable_needle_map() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = reload_as_tiered(dir, "vif_tierdown_test", 2);
assert!(matches!(v.nm, Some(NeedleMap::SortedFile(_))));
// The tail of the tier-down handler, once the .dat is back on disk.
v.update_remote_files(|files| {
files.remove(0);
})
.unwrap();
v.save_volume_info().unwrap();
v.open_local_dat_backend().unwrap();
assert!(!v.has_remote_file());
assert!(
!v.is_read_only(),
"tier-down should publish a writable volume"
);
assert!(
!matches!(v.nm, Some(NeedleMap::SortedFile(_))),
"tier-down must swap the read-only map out before writes are allowed"
);
let mut n = Needle {
id: NeedleId(7),
cookie: Cookie(0x5678),
data: b"after-tier-down".to_vec(),
data_size: 15,
..Needle::default()
};
v.write_needle(&mut n, true, true).unwrap();
v.sync_to_disk().unwrap();
drop(v);
// Reload: the needle must come back from .idx, not just have landed in
// .dat as unreferenced bytes.
let v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
let mut probe = Needle {
id: NeedleId(7),
..Needle::default()
};
v.read_needle(&mut probe).unwrap();
assert_eq!(probe.data, b"after-tier-down");
crate::remote_storage::s3_tier::global_s3_tier_registry()
.write()
.unwrap()
.remove("s3.vif_tierdown_test");
}
// An .idx that cannot be read up to the recorded size must abort
// compaction. Treating the short scan as the complete live set would
// commit a volume missing every needle past the truncation.
#[test]
fn test_compaction_aborts_on_truncated_index() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = make_test_volume(dir);
for i in 1..=4u64 {
write_test_needle(&mut v, i, format!("needle-{i}").as_bytes());
}
v.set_read_only_persist(false, true).unwrap();
v.sync_to_disk().unwrap();
}
let mut v = make_test_volume(dir);
let Some(NeedleMap::SortedFile(_)) = v.nm else {
panic!("read-only volume should search the on-disk .sdx");
};
let idx = OpenOptions::new()
.write(true)
.open(v.file_name(".idx"))
.unwrap();
idx.set_len(NEEDLE_MAP_ENTRY_SIZE as u64).unwrap();
drop(idx);
let err = v
.compact_by_index(0, 0, |_| true)
.expect_err("compaction must not commit a partially scanned index");
assert!(
matches!(err, VolumeError::Io(ref e) if e.kind() == io::ErrorKind::UnexpectedEof),
"unexpected error: {err:?}"
);
}
// Building .sdx writes to the index directory, which a read-only volume's
// may not allow. Before the sorted map that volume mounted and served reads
// off an in-memory index; it must still do so rather than going offline.
#[test]
#[cfg(unix)]
fn test_read_only_volume_mounts_when_the_index_dir_is_not_writable() {
use std::os::unix::fs::PermissionsExt;
// root ignores the directory mode, so there is nothing to simulate.
// SAFETY: `geteuid` takes no arguments, reads no memory and cannot fail.
if unsafe { libc::geteuid() } == 0 {
return;
}
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
{
let mut v = make_test_volume(&dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0x1234),
data: b"still-readable".to_vec(),
data_size: 14,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.set_read_only_persist(false, true).unwrap();
v.sync_to_disk().unwrap();
}
let set_mode = |mode: u32| {
let mut perms = std::fs::metadata(tmp.path()).unwrap().permissions();
perms.set_mode(mode);
std::fs::set_permissions(tmp.path(), perms).unwrap();
};
set_mode(0o555);
let loaded = Volume::new(
&dir,
&dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
);
set_mode(0o755); // restore before any assertion, so cleanup works
let v = loaded.expect("a read-only volume must mount without writing its index dir");
assert!(
!matches!(v.nm, Some(NeedleMap::SortedFile(_))),
"the .sdx could not be built, so the index should have fallen back to memory"
);
let mut probe = Needle {
id: NeedleId(1),
..Needle::default()
};
v.read_needle(&mut probe).unwrap();
assert_eq!(probe.data, b"still-readable");
}
// Compaction copies the .idx prefix index_file_size() reports. A read-only
// volume's map has no writer, so that must still be the loaded size, or the
// copy is empty and the commit drops every needle.
#[cfg(unix)]
fn check_read_only_compaction_keeps_needles(index_dir_writable: bool) {
use std::os::unix::fs::PermissionsExt;
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
{
let mut v = make_test_volume(&dir);
for i in 1..=3u64 {
write_test_needle(&mut v, i, format!("data-{i}").as_bytes());
}
v.set_read_only_persist(false, true).unwrap();
v.sync_to_disk().unwrap();
}
let idx_len = fs::metadata(format!("{dir}/1.idx")).unwrap().len();
let set_mode = |mode: u32| {
let mut perms = std::fs::metadata(tmp.path()).unwrap().permissions();
perms.set_mode(mode);
std::fs::set_permissions(tmp.path(), perms).unwrap();
};
if !index_dir_writable {
set_mode(0o555);
}
let loaded = Volume::new(
&dir,
&dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
);
set_mode(0o755);
let mut v = loaded.unwrap();
assert_eq!(
matches!(v.nm, Some(NeedleMap::SortedFile(_))),
index_dir_writable
);
assert_eq!(v.idx_file_size(), idx_len);
v.compact_by_index(0, 0, |_| true).unwrap();
v.commit_compact().unwrap();
for i in 1..=3u64 {
let mut n = Needle {
id: NeedleId(i),
..Needle::default()
};
v.read_needle(&mut n).unwrap();
assert_eq!(n.data, format!("data-{i}").as_bytes());
}
}
#[test]
#[cfg(unix)]
fn test_compacting_a_read_only_volume_keeps_its_needles_sorted_index() {
check_read_only_compaction_keeps_needles(true);
}
// The .sdx could not be built, so the index fell back to memory.
#[test]
#[cfg(unix)]
fn test_compacting_a_read_only_volume_keeps_its_needles_in_memory_fallback() {
// root ignores the directory mode, so there is nothing to simulate.
// SAFETY: `geteuid` takes no arguments, reads no memory and cannot fail.
if unsafe { libc::geteuid() } == 0 {
return;
}
check_read_only_compaction_keeps_needles(false);
}
// set_writable clears the read-only flags before it can know the .idx writer
// will attach. If attaching fails the flags have to go back: a volume that
// advertises writable while its map has no writer takes puts into memory and
// loses them at the next restart.
#[test]
#[cfg(unix)]
fn test_set_writable_rolls_back_when_the_idx_writer_cannot_attach() {
use std::os::unix::fs::PermissionsExt;
// root ignores the directory mode, so there is nothing to simulate.
// SAFETY: `geteuid` takes no arguments, reads no memory and cannot fail.
if unsafe { libc::geteuid() } == 0 {
return;
}
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
{
let mut v = make_test_volume(&dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0x1234),
data: b"before".to_vec(),
data_size: 6,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.set_read_only_persist(false, true).unwrap();
v.sync_to_disk().unwrap();
}
// A read-only mount, both halves of it: nothing new can be created in the
// directory, so .sdx generation fails and the index falls back to memory,
// and the .idx itself cannot be opened for write, so no writer attaches.
let set_mode = |path: &std::path::Path, mode: u32| {
let mut perms = std::fs::metadata(path).unwrap().permissions();
perms.set_mode(mode);
std::fs::set_permissions(path, perms).unwrap();
};
let idx_path = tmp.path().join("1.idx");
set_mode(&idx_path, 0o444);
set_mode(tmp.path(), 0o555);
let loaded = Volume::new(
&dir,
&dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
);
let marked = loaded.map(|mut v| {
let r = v.set_writable();
(
v.no_write_or_delete,
v.is_read_only(),
v.nm.as_ref().unwrap().has_idx_writer(),
r,
)
});
// Restore before any assertion, so cleanup works.
set_mode(tmp.path(), 0o755);
set_mode(&idx_path, 0o644);
let (no_write_or_delete, read_only, has_writer, result) = marked.unwrap();
assert!(
result.is_err(),
"set_writable cannot succeed without an .idx writer"
);
assert!(!has_writer, "the writer is what failed to attach");
assert!(
no_write_or_delete && read_only,
"a failed mark must leave the volume read-only, not writable with an index it cannot append to"
);
}
// .idx is a run of fixed-size rows and writers append at EOF, so a partial
// row left by a short write puts every row appended after it off alignment
// and the next load parses the rest as garbage. The tail has to go before
// anything writable attaches.
#[test]
fn test_writable_load_trims_a_torn_idx_tail() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
{
let mut v = make_test_volume(&dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0x1234),
data: b"first".to_vec(),
data_size: 5,
..Needle::default()
};
v.write_needle(&mut n, true, true).unwrap();
v.sync_to_disk().unwrap();
}
let idx_path = tmp.path().join("1.idx");
let whole = std::fs::metadata(&idx_path).unwrap().len();
{
use std::io::Write as _;
let mut f = OpenOptions::new().append(true).open(&idx_path).unwrap();
f.write_all(&[0xcd; 5]).unwrap();
}
let mut v = Volume::new(
&dir,
&dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert_eq!(
std::fs::metadata(&idx_path).unwrap().len(),
whole,
"the partial row should be gone before any writer attaches"
);
let mut n = Needle {
id: NeedleId(2),
cookie: Cookie(0x5678),
data: b"second".to_vec(),
data_size: 6,
..Needle::default()
};
v.write_needle(&mut n, true, true).unwrap();
v.sync_to_disk().unwrap();
drop(v);
// Both rows have to survive the round trip; a misaligned append loses
// the second one and garbles whatever follows it.
let v = Volume::new(
&dir,
&dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
for (id, want) in [(1u64, &b"first"[..]), (2u64, &b"second"[..])] {
let mut probe = Needle {
id: NeedleId(id),
..Needle::default()
};
v.read_needle(&mut probe).unwrap();
assert_eq!(probe.data, want);
}
}
// A delete-capable read-only volume — every tiered one — takes the in-memory
// loader's read-write branch, which fails on the same unwritable directory
// that refused the .sdx. It has to end up read-only rather than offline.
#[test]
#[cfg(unix)]
fn test_delete_capable_volume_falls_all_the_way_back_to_read_only() {
use std::os::unix::fs::PermissionsExt;
// root ignores the directory mode, so there is nothing to simulate.
// SAFETY: `geteuid` takes no arguments, reads no memory and cannot fail.
if unsafe { libc::geteuid() } == 0 {
return;
}
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
{
let mut v = make_test_volume(&dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0x1234),
data: b"still-readable".to_vec(),
data_size: 14,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.set_read_only_persist(true, true).unwrap();
v.sync_to_disk().unwrap();
}
let set_mode = |path: &std::path::Path, mode: u32| {
let mut perms = std::fs::metadata(path).unwrap().permissions();
perms.set_mode(mode);
std::fs::set_permissions(path, perms).unwrap();
};
let idx_path = tmp.path().join("1.idx");
set_mode(&idx_path, 0o444);
set_mode(tmp.path(), 0o555);
let loaded = Volume::new(
&dir,
&dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
);
set_mode(tmp.path(), 0o755);
set_mode(&idx_path, 0o644);
let v = loaded.expect("a delete-capable volume must mount without a writable index dir");
assert!(
v.no_write_or_delete,
"with no writer for tombstones, deletes have to be refused outright"
);
let mut probe = Needle {
id: NeedleId(1),
..Needle::default()
};
v.read_needle(&mut probe).unwrap();
assert_eq!(probe.data, b"still-readable");
}
// Issue #10937: a volume server holding hundreds of thousands of cloud-tiered
// volumes ran out of descriptors because every one of them pinned its index
// for the life of the process. A tiered volume must search the on-disk .sdx
// and hold nothing open while it sits idle.
#[test]
fn test_remote_volume_index_is_on_disk_and_holds_no_descriptors() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = make_test_volume(dir);
for i in 1..=4u64 {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(0x1234),
data: format!("needle-{i}").into_bytes(),
data_size: format!("needle-{i}").len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
// A writable volume does pin its .idx, which is also what proves the
// descriptor probe below is looking at something real.
if let Some(count) = crate::storage::needle_map::file_pool::open_index_fds(tmp.path()) {
assert!(count > 0, "a writable volume should hold its .idx open");
}
let vif = VifVolumeInfo {
files: vec![VifRemoteFile {
backend_type: "s3".to_string(),
backend_id: "vif_fd_test".to_string(),
key: "remote-key".to_string(),
offset: 0,
file_size: v.dat_file_size().unwrap(),
modified_time: 123,
extension: ".dat".to_string(),
}],
version: v.version().0 as u32,
..VifVolumeInfo::default()
};
std::fs::write(
format!("{}/1.vif", dir),
serde_json::to_string_pretty(&vif).unwrap(),
)
.unwrap();
let tier_config = crate::remote_storage::s3_tier::S3TierConfig {
access_key: "access".to_string(),
secret_key: "secret".to_string(),
region: "us-east-1".to_string(),
bucket: "bucket-a".to_string(),
endpoint: "http://127.0.0.1:1".to_string(),
storage_class: "STANDARD".to_string(),
force_path_style: true,
};
crate::remote_storage::s3_tier::global_s3_tier_registry()
.write()
.unwrap()
.register(
"s3.vif_fd_test".to_string(),
crate::remote_storage::s3_tier::S3TierBackend::new(&tier_config),
);
}
let mut v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert!(v.has_remote_file());
let Some(NeedleMap::SortedFile(ref nm)) = v.nm else {
panic!(
"tiered volume should search the on-disk .sdx, got {:?}",
v.nm.is_some()
);
};
let idx_path = nm.index_file_name().to_string();
let sdx_path = nm.db_file_name().to_string();
assert!(Path::new(&sdx_path).exists(), "load should build the .sdx");
let drop_pooled = || {
crate::storage::needle_map::file_pool::pooled_index_files().discard(&idx_path);
crate::storage::needle_map::file_pool::pooled_index_files().discard(&sdx_path);
};
if crate::storage::needle_map::file_pool::open_index_fds(tmp.path()).is_some() {
drop_pooled();
assert_eq!(
crate::storage::needle_map::file_pool::open_index_fds(tmp.path()),
Some(0),
"a loaded tiered volume must hold no .idx/.sdx descriptor"
);
}
// Lookups still resolve, straight off .sdx.
assert!(
!v.nm
.as_ref()
.unwrap()
.get(NeedleId(3))
.unwrap()
.unwrap()
.size
.is_deleted()
);
assert!(v.nm.as_ref().unwrap().get(NeedleId(9)).unwrap().is_none());
// Deletes are still allowed on a tiered volume and land in .idx.
let idx_before = std::fs::metadata(&idx_path).unwrap().len();
let deleted = v
.delete_needle(&mut Needle {
id: NeedleId(3),
cookie: Cookie(0x1234),
..Needle::default()
})
.unwrap();
assert!(deleted.0 > 0);
assert_eq!(
std::fs::metadata(&idx_path).unwrap().len(),
idx_before + NEEDLE_MAP_ENTRY_SIZE as u64,
"the tombstone must be appended to the .idx tail"
);
assert!(
v.nm.as_ref()
.unwrap()
.get(NeedleId(3))
.unwrap()
.unwrap()
.size
.is_deleted()
);
if crate::storage::needle_map::file_pool::open_index_fds(tmp.path()).is_some() {
drop_pooled();
assert_eq!(
crate::storage::needle_map::file_pool::open_index_fds(tmp.path()),
Some(0),
"reads and deletes must give their borrowed handles back"
);
}
crate::remote_storage::s3_tier::global_s3_tier_registry()
.write()
.unwrap()
.remove("s3.vif_fd_test");
}
#[test]
fn test_set_writable_keeps_remote_delete_only_mode() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let file_size = v.dat_file_size().unwrap();
v.update_remote_files(|files| {
files.push(PbRemoteFile {
backend_type: "s3".to_string(),
backend_id: "default".to_string(),
key: "remote-key".to_string(),
offset: 0,
file_size,
modified_time: 123,
extension: ".dat".to_string(),
})
})
.unwrap();
v.set_writable().unwrap();
assert!(v.is_read_only());
assert!(!v.no_write_or_delete);
assert!(v.no_write_can_delete);
}
// A volume booted with .vif ReadOnly=true used to come back stuck —
// load_index_inmemory built the CompactNeedleMap without an .idx writer
// attached, and set_writable only flipped the flag and rewrote .vif.
// The next put silently skipped the .idx append, so the write landed in
// memory only and was lost on the next restart.
#[test]
fn test_set_writable_reattaches_idx_writer_after_persisted_readonly() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: b"initial".to_vec(),
data_size: 7,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.set_read_only_persist(false, true).unwrap();
v.sync_to_disk().unwrap();
}
let mut v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert!(
v.no_write_or_delete,
"reloaded volume should be read-only from .vif"
);
assert!(
matches!(v.nm, Some(NeedleMap::SortedFile(_))),
"read-only load should search the on-disk .sdx"
);
v.set_writable().unwrap();
assert!(!v.is_read_only());
assert!(
!matches!(v.nm, Some(NeedleMap::SortedFile(_))),
"set_writable must swap out the read-only map, whose put always fails"
);
assert!(
v.nm.as_ref().unwrap().has_idx_writer(),
"set_writable must reattach the .idx writer or post-restart writes vanish"
);
let mut n = Needle {
id: NeedleId(2),
cookie: Cookie(2),
data: b"after-mark-writable".to_vec(),
data_size: 19,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
// Reload one more time — the .idx must contain the post-mark-writable
// entry, not just have it in memory.
drop(v);
let v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
let mut probe = Needle {
id: NeedleId(2),
..Needle::default()
};
v.read_needle(&mut probe).unwrap();
assert_eq!(
std::str::from_utf8(&probe.data).unwrap(),
"after-mark-writable"
);
}
// readOnlyCanDelete rejects writes, keeps accepting deletes, and survives
// a restart.
#[test]
fn test_read_only_can_delete_persists_and_keeps_deletes_working() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = make_test_volume(dir);
for id in 1..=2u64 {
let mut n = Needle {
id: NeedleId(id),
cookie: Cookie(id as u32),
data: b"payload".to_vec(),
data_size: 7,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.set_read_only_persist(true, true).unwrap();
assert!(v.no_write_can_delete);
assert!(!v.no_write_or_delete);
assert!(v.is_read_only());
let err = v
.write_needle(
&mut Needle {
id: NeedleId(3),
cookie: Cookie(3),
data: b"blocked".to_vec(),
data_size: 7,
..Needle::default()
},
true,
false,
)
.unwrap_err();
assert!(matches!(err, VolumeError::ReadOnly(_)));
let deleted = v
.delete_needle(&mut Needle {
id: NeedleId(1),
cookie: Cookie(1),
..Needle::default()
})
.unwrap();
assert!(deleted.0 > 0, "deletes must still land");
v.sync_to_disk().unwrap();
}
let mut v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert!(v.no_write_can_delete, "state must survive a restart");
assert!(!v.no_write_or_delete);
assert!(
v.nm.as_ref().unwrap().has_idx_writer(),
"canDelete boot must attach an .idx writer for delete tombstones"
);
let err = v
.write_needle(
&mut Needle {
id: NeedleId(4),
cookie: Cookie(4),
data: b"blocked".to_vec(),
data_size: 7,
..Needle::default()
},
true,
false,
)
.unwrap_err();
assert!(matches!(err, VolumeError::ReadOnly(_)));
let deleted = v
.delete_needle(&mut Needle {
id: NeedleId(2),
cookie: Cookie(2),
..Needle::default()
})
.unwrap();
assert!(deleted.0 > 0, "deletes must still land after restart");
// downgrading to plain readonly clears the canDelete flag
v.set_read_only_persist(false, true).unwrap();
assert!(v.no_write_or_delete);
assert!(!v.no_write_can_delete);
// -writable clears the state again.
v.set_writable().unwrap();
assert!(!v.is_read_only());
let mut n = Needle {
id: NeedleId(5),
cookie: Cookie(5),
data: b"writable-again".to_vec(),
data_size: 14,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
// Upgrading a volume that booted plain persisted-readonly to canDelete
// must attach the missing .idx writer, or delete tombstones cannot append.
#[test]
fn test_read_only_can_delete_upgrade_after_readonly_boot() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: b"payload".to_vec(),
data_size: 7,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.set_read_only_persist(false, true).unwrap();
v.sync_to_disk().unwrap();
}
let mut v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert!(v.no_write_or_delete);
let err = v
.delete_needle(&mut Needle {
id: NeedleId(1),
cookie: Cookie(1),
..Needle::default()
})
.unwrap_err();
assert!(
matches!(err, VolumeError::ReadOnly(_)),
"plain readonly must reject deletes"
);
v.set_read_only_persist(true, true).unwrap();
assert!(!v.no_write_or_delete);
assert!(v.no_write_can_delete);
assert!(
v.nm.as_ref().unwrap().has_idx_writer(),
"upgrade to canDelete must attach the .idx writer"
);
let deleted = v
.delete_needle(&mut Needle {
id: NeedleId(1),
cookie: Cookie(1),
..Needle::default()
})
.unwrap();
assert!(deleted.0 > 0, "deletes must land once canDelete is set");
}
#[test]
fn test_load_vif_defaults_local_version_and_bytes_offset() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let _v = make_test_volume(dir);
let vif = VifVolumeInfo::default();
std::fs::write(
format!("{}/1.vif", dir),
serde_json::to_string_pretty(&vif).unwrap(),
)
.unwrap();
}
let v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert_eq!(v.volume_info.version, Version::current().0 as u32);
assert_eq!(v.volume_info.bytes_offset, OFFSET_SIZE as u32);
assert_eq!(v.version(), Version::current());
}
#[test]
fn test_version_superblock_overrides_vif_version() {
// Go behavior: after reading the superblock, volumeInfo.Version is set
// to SuperBlock.Version, overriding whatever was in the .vif file.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let _v = make_test_volume(dir);
// Write a .vif with version=2, but the .dat superblock is version=3
let vif = VifVolumeInfo {
version: VERSION_2.0 as u32,
bytes_offset: OFFSET_SIZE as u32,
..VifVolumeInfo::default()
};
std::fs::write(
format!("{}/1.vif", dir),
serde_json::to_string_pretty(&vif).unwrap(),
)
.unwrap();
}
let v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
// Superblock version (3) overrides the .vif version (2)
assert_eq!(v.volume_info.version, VERSION_3.0 as u32);
assert_eq!(v.version(), VERSION_3);
}
#[test]
fn test_load_vif_rejects_bytes_offset_mismatch() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let _v = make_test_volume(dir);
let vif = VifVolumeInfo {
version: Version::current().0 as u32,
bytes_offset: (OFFSET_SIZE as u32) + 1,
..VifVolumeInfo::default()
};
std::fs::write(
format!("{}/1.vif", dir),
serde_json::to_string_pretty(&vif).unwrap(),
)
.unwrap();
}
let result = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
);
match result {
Ok(_) => panic!("expected bytes_offset mismatch to fail"),
Err(VolumeError::Io(io_err)) => {
assert_eq!(io_err.kind(), io::ErrorKind::InvalidData);
assert!(io_err.to_string().contains("bytes_offset mismatch"));
}
Err(other) => panic!("unexpected error: {other:?}"),
}
}
#[test]
fn test_remote_only_volume_load_reads_from_tier_backend() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let dat_bytes = {
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(7),
cookie: Cookie(0x7788),
data: b"remote-only".to_vec(),
data_size: 11,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
std::fs::read(v.file_name(".dat")).unwrap()
};
let dat_path = format!("{}/1.dat", dir);
std::fs::remove_file(&dat_path).unwrap();
let (endpoint, shutdown_tx) = spawn_fake_s3_server(dat_bytes.clone());
// Use a test-specific backend_id to avoid racing with other tests
// that share the global registry. Never call clear() — only
// register/remove our own entries.
let tier_config = crate::remote_storage::s3_tier::S3TierConfig {
access_key: "access".to_string(),
secret_key: "secret".to_string(),
region: "us-east-1".to_string(),
bucket: "bucket-a".to_string(),
endpoint,
storage_class: "STANDARD".to_string(),
force_path_style: true,
};
{
let mut registry = crate::remote_storage::s3_tier::global_s3_tier_registry()
.write()
.unwrap();
registry.register(
"s3.remote_only_rw".to_string(),
crate::remote_storage::s3_tier::S3TierBackend::new(&tier_config),
);
}
let vif = VifVolumeInfo {
files: vec![VifRemoteFile {
backend_type: "s3".to_string(),
backend_id: "remote_only_rw".to_string(),
key: "remote-key".to_string(),
offset: 0,
file_size: dat_bytes.len() as u64,
modified_time: 123,
extension: ".dat".to_string(),
}],
version: Version::current().0 as u32,
bytes_offset: OFFSET_SIZE as u32,
dat_file_size: dat_bytes.len() as i64,
..VifVolumeInfo::default()
};
std::fs::write(
format!("{}/1.vif", dir),
serde_json::to_string_pretty(&vif).unwrap(),
)
.unwrap();
let v = Volume::new(
dir,
dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
assert!(v.has_remote_file());
assert!(v.dat_file.is_none());
assert!(v.remote_dat_file.is_some());
let mut n = Needle {
id: NeedleId(7),
..Needle::default()
};
let size = v.read_needle(&mut n).unwrap();
assert_eq!(size, 11);
assert_eq!(n.data, b"remote-only");
let mut meta = Needle {
id: NeedleId(7),
..Needle::default()
};
let mut plan = v.needle_read_plan(NeedleId(7), false).unwrap();
plan.read_meta(&mut meta).unwrap();
let info = plan.into_stream_info(&meta);
assert!(matches!(info.source, NeedleStreamSource::Remote(_)));
let mut streamed = vec![0u8; info.data_size as usize];
info.source
.read_exact_at(&mut streamed, info.data_file_offset)
.unwrap();
assert_eq!(streamed, b"remote-only");
assert_eq!(meta.data_size, 11);
let _ = shutdown_tx.send(());
crate::remote_storage::s3_tier::global_s3_tier_registry()
.write()
.unwrap()
.remove("s3.remote_only_rw");
}
/// Volume destroy removes .vif alongside the primary data files.
#[test]
fn test_destroy_removes_vif() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: b"test".to_vec(),
data_size: 4,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
// Write a .vif file (as EC encode would)
let vif_path = format!("{}/1.vif", dir);
std::fs::write(&vif_path, r#"{"version":3}"#).unwrap();
assert!(std::path::Path::new(&vif_path).exists());
// .dat and .idx should exist
let dat_path = format!("{}/1.dat", dir);
let idx_path = format!("{}/1.idx", dir);
assert!(std::path::Path::new(&dat_path).exists());
assert!(std::path::Path::new(&idx_path).exists());
// Destroy the volume
v.destroy(false, false, false).unwrap();
// .dat and .idx should be gone
assert!(
!std::path::Path::new(&dat_path).exists(),
".dat should be removed"
);
assert!(
!std::path::Path::new(&idx_path).exists(),
".idx should be removed"
);
assert!(
!std::path::Path::new(&vif_path).exists(),
".vif should be removed"
);
}
/// Volume destroy with separate idx directory must clean up both dirs and .vif.
#[test]
fn test_destroy_with_separate_idx_dir() {
let dat_tmp = TempDir::new().unwrap();
let idx_tmp = TempDir::new().unwrap();
let dat_dir = dat_tmp.path().to_str().unwrap();
let idx_dir = idx_tmp.path().to_str().unwrap();
let mut v = Volume::new(
dat_dir,
idx_dir,
VolumeId(1),
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: b"hello".to_vec(),
data_size: 5,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
// Write .vif in data dir (as EC encode would)
let vif_path = format!("{}/1.vif", dat_dir);
std::fs::write(&vif_path, r#"{"version":3}"#).unwrap();
let dat_path = format!("{}/1.dat", dat_dir);
let idx_path = format!("{}/1.idx", idx_dir);
assert!(std::path::Path::new(&dat_path).exists());
assert!(std::path::Path::new(&idx_path).exists());
v.destroy(false, false, false).unwrap();
assert!(
!std::path::Path::new(&dat_path).exists(),
".dat removed from data dir"
);
assert!(
!std::path::Path::new(&idx_path).exists(),
".idx removed from idx dir"
);
assert!(
!std::path::Path::new(&vif_path).exists(),
".vif removed from data dir"
);
}
/// When an EC volume for the same id has a sealed .ecx on the same disk, the
/// .vif is shared with it and must survive the regular volume's deletion.
#[test]
fn test_destroy_keeps_vif_when_ec_coexists() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: b"test".to_vec(),
data_size: 4,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
let vif_path = format!("{}/1.vif", dir);
std::fs::write(&vif_path, r#"{"version":3}"#).unwrap();
// A sealed .ecx marks a coexisting EC volume for the same id.
let ecx_path = format!("{}/1.ecx", dir);
std::fs::write(&ecx_path, b"ec-index").unwrap();
v.destroy(false, false, false).unwrap();
let dat_path = format!("{}/1.dat", dir);
let idx_path = format!("{}/1.idx", dir);
assert!(!std::path::Path::new(&dat_path).exists(), ".dat removed");
assert!(!std::path::Path::new(&idx_path).exists(), ".idx removed");
assert!(
std::path::Path::new(&vif_path).exists(),
".vif kept: shared with the coexisting EC volume"
);
assert!(
std::path::Path::new(&ecx_path).exists(),
".ecx is an EC sidecar, never touched here"
);
}
/// A crash after the .idx/.dat were consumed but before the .cpx->.idx
/// rename committed leaves only .cpd + .cpx + .cpc. reconcile_compact_state
/// must roll the swap FORWARD and produce a loadable, writable volume
/// holding the compacted needle count.
#[test]
fn test_reconcile_roll_forward_marker_only() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for i in 1..=6u64 {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: format!("data-{}", i).into_bytes(),
data_size: format!("data-{}", i).len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
for id in [2u64, 5u64] {
let mut del = Needle {
id: NeedleId(id),
cookie: Cookie(id as u32),
..Needle::default()
};
v.delete_needle(&mut del).unwrap();
}
let expected_live = v.file_count() - v.deleted_count();
v.compact_by_index(0, 0, |_| true).unwrap();
v.close();
let cpd = v.file_name(".cpd");
let cpx = v.file_name(".cpx");
let cpc = v.file_name(".cpc");
let dat = v.file_name(".dat");
let idx = v.file_name(".idx");
assert!(Path::new(&cpd).exists());
assert!(Path::new(&cpx).exists());
// Simulate the mid-commit crash: original .dat/.idx are gone and only
// the compacted temp files plus the commit marker survive.
fs::remove_file(&dat).unwrap();
fs::remove_file(&idx).unwrap();
let _ = fs::remove_dir_all(v.file_name(".ldb"));
fs::write(&cpc, b"").unwrap();
v.reconcile_compact_state().unwrap();
assert!(!Path::new(&cpd).exists(), ".cpd consumed by roll-forward");
assert!(!Path::new(&cpx).exists(), ".cpx consumed by roll-forward");
assert!(!Path::new(&cpc).exists(), ".cpc cleared after swap");
assert!(Path::new(&dat).exists(), ".dat restored by roll-forward");
assert!(Path::new(&idx).exists(), ".idx restored by roll-forward");
// The rolled-forward files must load as a writable volume with the
// compacted needle count.
let reloaded = make_test_volume(dir);
assert!(!reloaded.is_read_only(), "rolled-forward volume read-only");
assert_eq!(
reloaded.file_count(),
expected_live,
"rolled-forward volume needle count"
);
}
/// With no .cpc marker, leftover .cpd/.cpx are an abandoned generation;
/// reconcile must roll BACK by deleting them and leave the live .dat/.idx.
#[test]
fn test_reconcile_roll_back_no_marker() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for i in 1..=4u64 {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: format!("data-{}", i).into_bytes(),
data_size: format!("data-{}", i).len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.compact_by_index(0, 0, |_| true).unwrap();
let cpd = v.file_name(".cpd");
let cpx = v.file_name(".cpx");
let dat = v.file_name(".dat");
let idx = v.file_name(".idx");
assert!(Path::new(&cpd).exists());
assert!(Path::new(&cpx).exists());
// No .cpc marker exists.
v.reconcile_compact_state().unwrap();
assert!(!Path::new(&cpd).exists(), ".cpd rolled back");
assert!(!Path::new(&cpx).exists(), ".cpx rolled back");
assert!(Path::new(&dat).exists(), "live .dat kept");
assert!(Path::new(&idx).exists(), "live .idx kept");
}
/// apply_compact_swap must abort without touching the live .dat/.idx when
/// the .cpd/.cpx temp files are missing (a stale or duplicate commit).
#[test]
fn test_apply_compact_swap_missing_temp_files_preserves_live() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for i in 1..=3u64 {
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: format!("data-{}", i).into_bytes(),
data_size: format!("data-{}", i).len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
let dat = v.file_name(".dat");
let idx = v.file_name(".idx");
let dat_len_before = fs::metadata(&dat).unwrap().len();
// No .cpd/.cpx present: the swap must refuse.
assert!(
v.apply_compact_swap().is_err(),
"swap should error when .cpd/.cpx are missing"
);
assert!(Path::new(&dat).exists(), "live .dat preserved");
assert!(Path::new(&idx).exists(), "live .idx preserved");
assert_eq!(
fs::metadata(&dat).unwrap().len(),
dat_len_before,
"live .dat unchanged"
);
}
}