07da302da0 volume server: ec.decode verifies, cleans up and compacts like Go, off the runtime (#11547)
* volume server: ec.decode reads the .ecx from the index dir it was copied to

VolumeEcShardsCopy writes the .ecx/.ecj into the receiver's -dir.idx, so
with a split data/index dir the decode target has no .ecx beside its
shards. VolumeEcShardsToVolume sized the .dat from the right .ecx but
built the .idx from the data dir, failing with NotFound after the .dat
was already published. It now reads .ecx/.ecj from where the EC volume
opened them and writes the .idx beside the .dat, where Go leaves it.

The live-entry check and the .dat size also ignored deletions recorded
only in the .ecj, which Go folds into the .ecx (RebuildEcxFile) first:
a fully deleted volume was decoded instead of reported as having no live
entries, and deleted tail needles were copied into the .dat. Both now
treat journaled ids as deleted, without rewriting the sealed .ecx.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps the decoded volume writable and reads every .ecj

The rebuilt .idx copied a journaled tail needle's .ecx row verbatim after
the .dat was cut short before it, so the mount saw a row past EOF and
marked the decoded volume read-only. Rows of deleted needles the .dat no
longer holds are now dropped, and each journaled needle still in the .dat
gets one tombstone instead of one per journal entry.

VolumeEcShardsCopy appends journals collected from other holders into
the idx dir, but the decode read only the .ecj beside the .ecx, which
sits in the data dir when this server generated the shards. It now
reads both, once, in bounded chunks via the loader EcVolume uses.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: test ec.decode drops a sealed .ecx tail tombstone

Covers the other half of the rule added in the previous commit: a tail
needle tombstoned in the .ecx itself (Go's RebuildEcxFile) is cut from
the .dat, and its row must not reach the rebuilt .idx either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode runs its file I/O off the async runtime

VolumeEcShardsToVolume released the store lock before decoding, but read
the .ecx/.ecj, rebuilt the .dat and wrote the .idx inside the async
handler, parking a runtime worker for the length of a volume-sized copy.
The decode now runs in spawn_blocking on inputs snapshotted under the
store lock.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode checks the rebuilt .dat is complete

Go stats the decoded .dat before writing the .idx (VerifyDecodedDatFile)
and fails the decode when it is shorter than the extent the EC index
references, since the caller deletes the shards once the call returns.
The Rust handler returned success without that check. The rebuild
already fails on a short shard read, so this guards the published file
itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode drops the decoded volume's bitrot sidecars

Go removes <base>.ecsum and <base>.ecsum.v<N> beside the .dat and beside
the .ecx once the .idx is written, so a stale checksum sidecar cannot
pass for the protection of a later re-encode. The Rust handler left them
in place. Removal is best effort, as in Go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode compacts the decoded volume

Go ends VolumeEcShardsToVolume with an offline CompactVolumeFiles, so the
decoded volume holds only live needles. The Rust decode left every needle
deleted through the .ecj in the .dat, tombstoned in the .idx, until a
later vacuum reclaimed it.

Store::compact_volume_files loads the unmounted volume, checks free space
the way the vacuum does (the estimate now lives in one helper), and runs
the vacuum's compact-by-index and commit. As in Go a failed compaction is
logged and the decode still succeeds, so the uncompacted .idx rules stay:
the tests that pin them now make the compaction fail.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps deletes journaled while the .dat is written

The decode read the .ecj journals once, before rebuilding the .dat, so a
delete that reached the EC volume during the rebuild was left out of the
new .idx and the needle came back live. Each journal's read length is now
kept, and the bytes appended since are read just before the .idx is
written, after waiting out any journal append in flight (appends hold
the store write lock), so every delete acknowledged by then is in the
.idx. A delete after that point is still lost, as in Go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Guard overlapping ec decode requests; serialize journal catch-up

volume_ec_shards_to_volume runs its decode in spawn_blocking, so a
dropped request leaves the job running and a retry would race it on the
temporary and final volume files. Claim the vid in a per-server
in-flight set until the blocking job finishes, and return Unavailable
to an overlapping request. The Go handler has the same exposure and
gets the same guard.

Journal appends hold the store write lock through their
sync-or-truncate, so holding a read lock across the catch-up read
guarantees every record it sees is committed: a rolled-back delete can
no longer leave a tombstone in the decoded index.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Reconcile the swap when offline compaction commit fails

A CommitCompact that fails after the .cpc marker may have renamed .dat
but not .idx. cleanup_compact refuses while the marker exists, so the
mismatched pair survived until a restart reconciled it — and the decode
caller treats the failure as non-fatal. Run reconcileCompactState on
commit failure so a decided swap rolls forward and orphan temps are
removed before the volume can mount.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Release the decode claim on panic

* volume: add ec_decodes_in_flight to the integration-test state literal

* volume server: hold the decode tail's lock through compaction

The catch_up read released before the rebuilt .idx was written and the
volume compacted, so a delete synced to .ecj in that window was durably
journaled yet absent from the published index — resurrecting the needle.
Rust now holds the store read lock from catch_up through compact, and Go
mirrors it by holding the volume's journal lock from the journal-
consuming index write through CompactVolumeFiles.

* volume server: serialize ec decode's tail per volume, not per store

Review follow-ups on the decode path:

- Rust: holding the store read lock from journal catch-up through the
  offline compaction stalled every writer on unrelated volumes for the
  whole rewrite. The new ec_decode_tail set marks the vid only while its
  .idx is published and .cpd/.cpx swapped; the two local .ecj append paths
  (VolumeEcBlobDelete, the distributed delete's local journal) wait on a
  Notify for that span — Go's per-volume ecjFileAccessLock semantics
  without the global stall. VolumeMount and the staged-adopt path are also
  held off while a decode claim is in flight so neither can race the swap.

- Rust: the initial journal read ran unlocked, so bytes a rolled-back
  append later truncated could be folded in as phantom tombstones. The
  first pass stays unlocked (a slow journal must not stall the store) and
  a rescan under the quiescing read lock re-reads only committed content;
  catch_up now rebuilds the id set when a regular journal shrank.

- Go: the decode resolved the compaction DiskLocation through
  FindEcVolume while holding the journal lock, inverting DestroyEcVolume's
  map->journal order into a deadlock. The lookup now happens first, and
  DestroyEcVolume/deleteEcVolumeById/DiskLocation.Close destroy outside
  the map lock.

- Go: RebuildEcxFile unlinks .ecj while the volume's ecjFile handle stays
  open, so later deletes could commit to a detached inode. Both call sites
  now fold under the journal lock and ReopenDeletionJournal repoints the
  handle at the live path, working on the volume's resolved .ecx dir
  (EcIndexBaseFileName) rather than the configured index dir.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: fence EC remounts behind the destroy tombstone

DestroyEcVolume, deleteEcVolumeById, and the collection-delete sweep now
remove the EcVolume from ecVolumes before destroying it off-lock, so a
concurrent remount could re-open shard files that the in-flight destroy
then unlinks — registering a detached fd.

Each destroy records a per-vid tombstone channel in a new
ecVolumesDestroying map before dropping the map entry and closes it when
Destroy returns. The tombstone intentionally survives as the vid's
destroy generation: loadEcShardWithIdxDir compares it before and after
opening the shard, so a destroy that both started and finished inside the
open window is still detected. A mismatch drops the just-opened shard
(releasing its fd and mount gauge) and retries after the destroy
completes; a successful mount clears the stale tombstone.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: rescan the .ecj under the store lock only after a rollback

The decode's second journal pass ran a full rescan under the store read
lock on every decode, stalling unrelated writers for the length of the
scan. Bump a process-wide epoch whenever a failed append truncates its
uncommitted tail; an unchanged epoch between the unlocked read and the
quiesced pass proves every id folded in was committed, so catch_up()
suffices. catch_up() also treats a journal that was read but has since
disappeared as shrunk to zero, so its earlier ids cannot linger.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: check the decode tail under the store write lock on delete

A blob delete waited for the publishing tail before taking the store
write lock, so a decode that claimed the tail while the delete was
parked behind the decoder's read lock could still see the journal append
land after the rebuilt .idx — an acknowledged delete the mount would
miss. Test tail membership under the write lock instead, retrying after
the wait; journal_delete_local reports WouldBlock for the same recheck
on the distributed path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: claim the vid for mount and staged adoption, per volume

VolumeMount and the staged .copying adoption held the
ec_decodes_in_flight set lock through slow file renames and mounts,
stalling every unrelated volume's decode, mount, and adoption. Take the
per-volume claim instead — the same exclusion against a racing decode
for this vid, released when the call returns.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: fail the decode when a compaction commit marker survives

CompactVolumeFiles' caller logged a compaction error and went on to
delete the EC shards. When the commit marker (.cpc) is still on disk the
.dat/.idx swap was decided but could not be reconciled, so the mounted
pair may be mismatched — report the failure instead so the shards are
kept and the caller can retry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: gate the parked-delete test on the held write lock

The releaser thread and the spawned delete raced for the store write
lock; on a slow runner the delete could acquire it first and commit
before the tail was ever claimed, failing !delete.is_finished() on the
Windows unit-test job. Spawn the delete only after the thread reports
the lock held.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:55:15 +08:00
2026-10-03 00:54:20 +00:00
2019-04-30 03:23:20 +00:00
2023-01-05 11:01:22 -08:00

SeaweedFS

Slack Twitter Build Status GoDoc Wiki Docker Pulls SeaweedFS on Maven Central Artifact Hub

SeaweedFS Logo

SeaweedFS is a simple and highly scalable distributed file system. There are two objectives:

  1. to store billions of files!
  2. to serve the files fast!

One weed binary serves an S3 object store, a POSIX file system, and a lakehouse with S3 Tables, all over the same data. Each blob is one disk read away, capacity grows by starting another volume server, and cloud storage can be cached or tiered transparently. Both read and write operations have O(1) complexity and can run at the full speed supported by the underlying hardware.

Table of Contents

Quick Start

One command

Download the latest binary from the releases page and unzip the single weed (or weed.exe) file, or let the install script put it in /usr/local/bin:

curl -fsSL https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/install.sh | bash

Then start a ready-to-use S3 object store:

AWS_ACCESS_KEY_ID=admin \
AWS_SECRET_ACCESS_KEY=secret \
S3_BUCKET=my-bucket \
./weed mini -dir=./data

That's it. The S3 endpoint is at http://localhost:8333, my-bucket exists, and admin/secret are valid credentials:

AWS_ACCESS_KEY_ID=admin AWS_SECRET_ACCESS_KEY=secret \
  aws --endpoint-url http://localhost:8333 s3 cp README.md s3://my-bucket/

The same process also runs the master, a volume server, the filer, WebDAV, the Iceberg REST catalog, and the Admin UI. Add S3_TABLE_BUCKET=warehouse to also create an Iceberg table bucket, or warehouse:LANCE for a Lance one. Drop the AWS keys to run without authentication for development.

macOS: if the binary is quarantined, run xattr -d com.apple.quarantine ./weed first.

weed mini is auto-tuned for one node and is fine for single-node production, such as an S3 gateway that issues presigned URLs. See Quick Start with weed mini.

Docker

docker run -p 8333:8333 -v weed-data:/data \
  -e AWS_ACCESS_KEY_ID=admin \
  -e AWS_SECRET_ACCESS_KEY=secret \
  -e S3_BUCKET=my-bucket \
  chrislusf/seaweedfs

Same behavior as the weed mini command above.

Docker Compose

To run master, volume server, filer, S3, and WebDAV as separate services:

wget https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/docker/seaweedfs-compose.yml
wget -P prometheus https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/docker/prometheus/prometheus.yml
docker compose -f seaweedfs-compose.yml -p seaweedfs up

Docker Compose for S3 adds credentials, and the docker/compose folder has variants for replication, mounts, message queues, and more.

Kubernetes with Helm

helm repo add seaweedfs https://seaweedfs.github.io/seaweedfs/helm
helm install seaweedfs seaweedfs/seaweedfs -n seaweedfs --create-namespace -f values.yaml

A production-shaped values.yaml for a three-node cluster: two copies of every write, three masters, and an S3 endpoint with credentials and a bucket.

global:
  seaweedfs:
    enableReplication: true
    replicationPlacement: "001"   # one extra copy on another server; "002" for two

master:
  replicas: 3
  data:
    type: persistentVolumeClaim   # the cluster's default storage class; add storageClass to pick one
    size: 1Gi

volume:
  replicas: 3                     # at least 1 + the sum of the replication digits
  dataDirs:
    - name: data
      type: persistentVolumeClaim
      size: 500Gi
      maxVolumes: 0               # size the volume count from the disk

filer:
  replicas: 2
  data:
    type: persistentVolumeClaim
    size: 20Gi

s3:
  enabled: true
  replicas: 2
  enableAuth: true
  credentials:
    admin:
      accessKey: admin
      secretKey: change-me
  createBuckets:
    - name: app-storage

The S3 endpoint is the seaweedfs-s3 service on port 8333. Helm Chart Recipes has values for a development cluster, a lakehouse with the Iceberg catalog exposed, filer metadata on PostgreSQL, and node-local disks. The SeaweedFS Operator and the CSI driver are the other Kubernetes paths.

Build from source

git clone https://github.com/seaweedfs/seaweedfs.git
cd seaweedfs/weed && make install

weed lands in $GOPATH/bin. Getting Started covers running master, volume, filer, and S3 as separate processes.

Scale out

Capacity is a volume server. Start one on any machine with disk and point it at the master:

weed volume -dir=/data -master=<master_host>:9333

Nothing rebalances until you ask it to. Throughput is a filer or S3 gateway; they are stateless, so run as many as you need behind a load balancer. Production Setup walks through a multi-node cluster.

Back to TOC

Why SeaweedFS

Fast

  • One disk read per blob. A small file is one blob; a large file is split into chunks of a few MB, each its own blob. A volume server keeps a 16-byte index entry per blob in memory and reads it in a single seek, also for erasure-coded data.
  • The master is not in the read path. Clients cache the volume-to-server mapping and talk to volume servers directly.
  • 40 bytes of metadata per file on disk. Small files are packed into append-only volume files, so there is no per-file inode, no per-file metadata file, no fragmentation, and writes are SSD friendly.
  • Hot data is replicated; erasure coding is applied to warm data in the background, so writes never pay the encoding cost.
  • The Rust volume server is a drop-in for higher throughput and lower tail latency on the same on-disk format.

On one laptop, weed benchmark writes 1KB files at 15,700 per second and reads them back at 47,000 per second, and a mixed S3 warp run totals 3.2 GiB/s. Numbers are in the Benchmark section; throughput grows with volume servers and gateways.

Scalable

  • The master tracks volumes, not files. A cluster with billions of files has a few thousand volumes, so the master stays small. One master is enough for most clusters; run three for Raft failover.
  • Adding a server adds capacity with no data reshuffle. Balancing, vacuum, erasure coding, and repair run on demand from weed shell or the maintenance worker.
  • Filer and S3 gateways are stateless and scale linearly. Directory metadata lives in a store you already run: LevelDB, RocksDB, SQLite, MySQL, PostgreSQL, Cassandra, HBase, MongoDB, Redis, Elasticsearch, etcd, TiKV, FoundationDB, YDB, ArangoDB, Tarantool, and MySQL or PostgreSQL compatible databases such as TiDB, CockroachDB, and MemSQL.
  • Rack and data center aware replication, tiered storage across disk types, and transparent cloud tiering for unlimited capacity.
  • Files from a byte to tens of TB. Volumes up to 8TB with the large-disk build.

The most complete S3 API

The S3 gateway implements the object, bucket, S3 Tables, IAM, and STS APIs on one endpoint, so the AWS SDKs and CLI, rclone, restic, Spark, and Trino work unchanged.

API Operations
S3 bucket and object 73
S3 Tables 36
IAM 39
STS 5

The full operation list is in Amazon S3 API, and Supported APIs vs MinIO compares. The S3 compatibility suite and the SDK, IAM, SSE, policy, and Spark integration tests run in CI on every change.

A data warehouse with S3 Tables

SeaweedFS is a lakehouse in one system. S3 Table Buckets hold Apache Iceberg tables by default, or Lance tables for vectors and multimodal data, and the built-in Iceberg REST Catalog and Lance namespace serve them directly. There is no Hive Metastore, Glue, or separate catalog service to deploy, secure, and back up.

S3_TABLE_BUCKET=warehouse ./weed mini -dir=./data brings the whole stack up on a laptop.

A fast cache for cloud storage

Cloud Drive mounts a bucket from S3, Google Cloud Storage, Azure, Backblaze B2, Wasabi, Storj, or any S3-compatible store into SeaweedFS and serves it at local speed:

  • Metadata is pulled once, so listing, stat, and directory walks cost no cloud API calls.
  • File content is downloaded once, on first read or warmed by folder, name pattern, size, or age, and cached with the capacity of the whole cluster: cache everything, no churn.
  • Local writes complete at local latency and are written back to the cloud asynchronously in the cloud's native layout, so other tools keep reading the bucket directly.
  • Uncache by the same rules to free local disk while keeping the metadata.

Cloud Tier goes the other direction, moving whole warm volumes to cloud storage while keeping one-read access, and the Gateway to Remote Object Storage mirrors every bucket to a remote store. Faster and cheaper than reading the cloud directly.

Active-active replication and more

Back to TOC

Architecture

SeaweedFS Architecture

  • Master servers, one or a Raft group of three, track which volume lives on which volume server and hand out file ids. They are not in the read path.
  • Volume servers store blobs in append-only volume files, keep a 16-byte in-memory index per blob, and replicate or erasure-code at the volume level.
  • Filer servers add directories and files on top, with metadata in a store of your choice, and expose HTTP, S3, WebDAV, SFTP, FUSE, and the table catalogs.

The blob store started from Facebook's Haystack, erasure coding takes ideas from f4, and the whole has a lot in common with Tectonic and Colossus. How file ids are assigned, written, and looked up, and why a master that tracks volumes scales, is in Blob Store Architecture; the services are in Components and the white paper.

Back to TOC

Compared to Other Systems

Most other distributed file systems seem more complicated than necessary.

SeaweedFS is meant to be fast and simple, in both setup and operation. If you do not understand how it works when you reach here, we've failed! Please raise an issue with any questions or update this file with clarifications.

SeaweedFS is constantly moving forward. Same with other systems. These comparisons can be outdated quickly. Please help to keep them updated.

Compared to HDFS

HDFS uses the chunk approach for each file, and is ideal for storing large files.

SeaweedFS is ideal for serving relatively smaller files quickly and concurrently.

SeaweedFS can also store extra large files by splitting them into manageable data chunks, and store the file ids of the data chunks into a meta chunk. This is managed by "weed upload/download" tool, and the weed master or volume servers are agnostic about it.

Compared to GlusterFS, Ceph

The architectures are mostly the same. SeaweedFS aims to store and read files fast, with a simple and flat architecture. The main differences are

  • SeaweedFS optimizes for small files, ensuring O(1) disk seek operation, and can also handle large files.
  • SeaweedFS statically assigns a volume id for a file. Locating file content becomes just a lookup of the volume id, which can be easily cached.
  • SeaweedFS Filer metadata store can be any well-known and proven data store, e.g., Redis, Cassandra, HBase, Mongodb, Elastic Search, MySql, Postgres, Sqlite, MemSql, TiDB, CockroachDB, Etcd, YDB etc, and is easy to customize.
  • SeaweedFS Volume server also communicates directly with clients via HTTP, supporting range queries, direct uploads, etc.
System File Metadata File Content Read POSIX REST API Optimized for large number of small files
SeaweedFS lookup volume id, cacheable O(1) disk seek Yes Yes
SeaweedFS Filer Linearly Scalable, Customizable O(1) disk seek FUSE Yes Yes
GlusterFS hashing FUSE, NFS
Ceph hashing + rules FUSE Yes
MooseFS in memory FUSE No
MinIO separate meta file per drive for each file Yes No
RustFS separate meta file per drive for each file Yes No

GlusterFS stores files, both directories and content, in configurable volumes called "bricks". It hashes the path and filename into ids, and assigned to virtual volumes, and then mapped to "bricks".

Compared to MooseFS

MooseFS chooses to neglect small file issue. From moosefs 3.0 manual, "even a small file will occupy 64KiB plus additionally 4KiB of checksums and 1KiB for the header", because it "was initially designed for keeping large amounts (like several thousands) of very big files"

MooseFS Master Server keeps all meta data in memory. Same issue as HDFS namenode.

Compared to Ceph

Ceph can be setup similar to SeaweedFS as a key->blob store. It is much more complicated, with the need to support layers on top of it. Here is a more detailed comparison

SeaweedFS has a centralized master group to look up free volumes, while Ceph uses hashing and metadata servers to locate its objects. Having a centralized master makes it easy to code and manage.

Ceph, like SeaweedFS, is based on the object store RADOS. Ceph is rather complicated with mixed reviews.

Ceph uses CRUSH hashing to automatically manage data placement, which is efficient to locate the data. But the data has to be placed according to the CRUSH algorithm. Any wrong configuration would cause data loss. Topology changes, such as adding new servers to increase capacity, will cause data migration with high IO cost to fit the CRUSH algorithm. SeaweedFS places data by assigning them to any writable volumes. If writes to one volume failed, just pick another volume to write. Adding more volumes is also as simple as it can be.

SeaweedFS is optimized for small files. Small files are stored as one continuous block of content, with at most 8 unused bytes between files. Small file access is O(1) disk read.

SeaweedFS Filer uses off-the-shelf stores, such as MySql, Postgres, Sqlite, Mongodb, Redis, Elastic Search, Cassandra, HBase, MemSql, TiDB, CockroachCB, Etcd, YDB, to manage file directories. These stores are proven, scalable, and easier to manage.

SeaweedFS comparable to Ceph advantage
Master MDS simpler
Volume OSD optimized for small files
Filer Ceph FS linearly scalable, Customizable, O(1) or O(logN)

Compared to MinIO, RustFS

Please note, as Apr 25, 2026 MinIO ceased development. It's strongly discouraged to use that unmaintained software with multiple security bugs. RustFS is a MinIO reimplementation in Rust, Apache 2.0 licensed and still developed, keeping MinIO's storage model down to a byte-compatible on-disk format. So the points below apply to both.

MinIO followed AWS S3 closely and was ideal for testing for S3 API. It had good UI, policies, versionings, etc. SeaweedFS is trying to catch up here.

The metadata are in simple files. Each file write incurs extra writes to the corresponding meta file, on every drive of the erasure set. Changing only tags or retention rewrites that meta file on all of them, so the write amplification does not shrink with object size.

There is no optimization for lots of small files. The files are simply stored as is to local disks. Plus the extra meta file and shards for erasure coding, it only amplifies the LOSF problem.

Multiple disk IO are needed to read one file. SeaweedFS has O(1) disk reads, even for erasure coded files.

Erasure coding is full-time. SeaweedFS uses replication on hot data for faster speed and optionally applies erasure coding on warm data.

No POSIX-like API support.

There are specific requirements on storage layout, which makes it hard to scale out and to maintain. An erasure set must be 2 to 16 drives and must divide the drive list symmetrically, and capacity grows or shrinks a whole pool at a time. In SeaweedFS, just start one volume server pointing to the master. That's all.

Back to TOC

Benchmark

Unscientific single-machine numbers from a MacBook with an SSD. weed benchmark, 1 million 1KB files, concurrency 16:

Requests per second p50 p99
Write 15,708 0.8 ms 2.6 ms
Random read 47,019 0.3 ms 0.7 ms

make benchmark runs warp mixed S3 traffic against a local weed server:

Mixed operations.
Operation: DELETE, 10%, Concurrency: 20, Ran 42s.
 * Throughput: 55.13 obj/s

Operation: GET, 45%, Concurrency: 20, Ran 42s.
 * Throughput: 2477.45 MiB/s, 247.75 obj/s

Operation: PUT, 15%, Concurrency: 20, Ran 42s.
 * Throughput: 825.85 MiB/s, 82.59 obj/s

Operation: STAT, 30%, Concurrency: 20, Ran 42s.
 * Throughput: 165.27 obj/s

Cluster Total: 3302.88 MiB/s, 550.51 obj/s over 43s.

Read throughput is bounded by the random read speed of the disks, and grows with every volume server added. More numbers, including multi-node, FUSE, and Hadoop, are in Benchmarks, S3 API Benchmark, FIO benchmark, and Independent Benchmarks.

Back to TOC

Enterprise

For enterprise users, please visit seaweedfs.com for the SeaweedFS Enterprise Edition, which has advanced features, including data recovery, self-healing storage, customizable erasure coding, EC vacuum and repair, etc.

Back to TOC

License

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.

Back to TOC

Sponsors

Sponsor SeaweedFS via Patreon

SeaweedFS is an independent Apache-licensed open source project with its ongoing development made possible entirely thanks to the support of these awesome backers. If you'd like to grow SeaweedFS even stronger, please consider joining our sponsors on Patreon.

Your support will be really appreciated by me and other supporters!

Gold Sponsors

nodion piknik keepsec zyner

Back to TOC

Star History

Star History

Languages
Go 82.3%
Rust 9.5%
templ 2.9%
Java 1.8%
Shell 0.9%
Other 2.4%