filer: option to store system metadata logs in their own collection (#11551)

* feat(filer): option to store system metadata logs in their own collection

The filer's internal /topics/.system/log chunks are assigned to the
filer's default collection (-collection). In a multi-filer deployment
that default is often empty, so every restart flap, full-sync, or
event-buffered flush grows the default collection with system chunks that
are indistinguishable from user data in collection.list. This is a large
part of what makes the default collection balloon and confuses orphan
analysis.

This keeps the internal log in a dedicated collection when the operator
asks for one, without changing where user data goes:

- New optional override, filer.options.metaLog.collection (and
  .replication), read in NewFiler so both `weed filer` and
  `weed server -filer` honour it. Default "" => exactly today's
  behaviour (log follows the filer default), fully backward compatible.
- Resolution is a small helper: override first, then the filer default,
  then a storage rule matched on the log path. Kept separate from the
  user write path so the internal log targets itself.
- bucketCollection() is hardened the same way it already protects the
  filer's default collection: a bucket that happens to resolve to the
  redirected meta-log collection must not drop it on delete, because it
  backs internal log volumes.
- Scaffold filer.toml documents the new knobs under [filer.options].

Related to the persisted deletion ledger branch (fix/persist-deletion-queue):
together they cut the two sources of post-flap junk in the default
collection — that PR stops orphaned user-chunk leak on filer crash,
this one stops the internal log from living in default at all. They are
independent: no file overlap, no functional dependency; either can merge
first. They are paired only in the narrative of cleaning up default.

Adds unit tests for the collection/replication resolution chain, the
viper keys, and the bucket-delete guard (run green under -race).

Co-Authored-By: Athena 🏛️ <hermes-agent@local> (custom / Qwen3.8-Flash-Next-ROCmFP4)

* filer: collect bucket chunks when its collection survives the delete

bucketCollection returning "" preserves the collection, but the bucket
path still skipped per-entry chunk collection and could skip listing the
children entirely, so a bucket sharing the meta-log (or any preserved)
collection left its object chunks orphaned with no entry pointing at
them. Only the wholesale drop of a deleted collection skips those now.

Note in filer.toml that the meta-log target should stay stable: chunks
written under an older collection are not migrated.

* filer: exercise the metaLog override wiring through NewFiler

The viper test only echoed back the keys it set, so a wrong key in
NewFiler would still pass. It now asserts the fields NewFiler fills
from those keys.

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: tighten comments around the metaLog collection override

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
This commit is contained in:
authored and GitHub committed 2026-10-02 08:50:24 +08:00
1 parent d8f926cf46
commit 8d97284d0a
5 files changed
+199 -16

No files matched your search

+7
View File
@@ -16,6 +16,13 @@ recursive_delete = false
# for S3: how long to wait before deleting an empty folder.
# increase this if using tools like Spark that create temporary directories.
#s3.empty_folder_cleanup_delay = "2m"
# where the filer's internal system metadata log chunks are stored.
# By default the log follows the filer's -collection flag; set this to keep
# the internal log in its own collection, out of the default one.
# Keep the target stable: log chunks written under an older value are not
# moved, and old log entries keep referencing that collection's volumes.
#metaLog.collection = "filer-meta"
#metaLog.replication = ""
####################################################
# The following are filer store options
+29 -11
View File
@@ -46,17 +46,22 @@ var (
)
type Filer struct {
UniqueFilerId int32
UniqueFilerEpoch int32
Store VirtualFilerStore
MasterClient *wdclient.MasterClient
FileIdDeletionQueue *util.UnboundedQueue
GrpcDialOption grpc.DialOption
DirBucketsPath string
Cipher bool
LocalMetaLogBuffer *log_buffer.LogBuffer
metaLogCollection string
metaLogReplication string
UniqueFilerId int32
UniqueFilerEpoch int32
Store VirtualFilerStore
MasterClient *wdclient.MasterClient
FileIdDeletionQueue *util.UnboundedQueue
GrpcDialOption grpc.DialOption
DirBucketsPath string
Cipher bool
LocalMetaLogBuffer *log_buffer.LogBuffer
metaLogCollection string
metaLogReplication string
// Override where system metadata-log chunks are assigned, keeping the
// internal log out of the default collection; empty keeps today's
// behaviour. Set via viper: filer.options.metaLog.collection / .replication.
metaLogTargetCollection string
metaLogTargetReplication string
DefaultDiskType string
MetaAggregator *MetaAggregator
Signature int32
@@ -111,6 +116,19 @@ func NewFiler(masters pb.ServerDiscovery, grpcDialOption grpc.DialOption, filerH
f.metaLogCollection = collection
f.metaLogReplication = replication
// Optional override for where the system metadata-log chunks land, so
// operators can keep internal log volumes out of the default collection.
// Unset (""), this changes nothing: meta logs keep following the filer
// default exactly as before.
v := util.GetViper()
v.SetDefault("filer.options.metaLog.collection", "")
v.SetDefault("filer.options.metaLog.replication", "")
f.metaLogTargetCollection = v.GetString("filer.options.metaLog.collection")
f.metaLogTargetReplication = v.GetString("filer.options.metaLog.replication")
if f.metaLogTargetCollection != "" {
glog.V(0).Infof("system metadata logs will be stored in collection %q", f.metaLogTargetCollection)
}
if newPlacementOverlay != nil {
f.placementOverlay = newPlacementOverlay(f)
}
+13 -3
View File
@@ -72,9 +72,12 @@ func (f *Filer) DeleteEntryMetaAndData(ctx context.Context, p util.FullPath, isR
if isDeleteCollection {
collectionName = f.bucketCollection(ctx, entry.Name())
}
// A preserved collection outlives the bucket, so its chunks are collected
// per entry rather than dropped wholesale with it.
dropsCollection := isDeleteCollection && collectionName != ""
if entry.IsDirectory() {
// delete the folder children, not including the folder itself
err = f.doBatchDeleteFolderMetaAndData(ctx, entry, isRecursive, ignoreRecursiveError, shouldDeleteChunks && !isDeleteCollection, isDeleteCollection, isFromOtherCluster, signatures, func(hardLinkIds []HardLinkId) error {
err = f.doBatchDeleteFolderMetaAndData(ctx, entry, isRecursive, ignoreRecursiveError, shouldDeleteChunks && !dropsCollection, isDeleteCollection && (dropsCollection || !shouldDeleteChunks), isFromOtherCluster, signatures, func(hardLinkIds []HardLinkId) error {
// A case not handled:
// what if the chunk is in a different collection?
if shouldDeleteChunks {
@@ -97,7 +100,7 @@ func (f *Filer) DeleteEntryMetaAndData(ctx context.Context, p util.FullPath, isR
return fmt.Errorf("delete file %s: %v", p, err)
}
if shouldDeleteChunks && !isDeleteCollection {
if shouldDeleteChunks && !dropsCollection {
if len(entry.HardLinkId) != 0 && entry.HardLinkCounter > 1 {
// if the file is a hard link and there are other hard links, do not delete the chunks
} else {
@@ -261,10 +264,17 @@ func (f *Filer) bucketCollection(ctx context.Context, bucket string) (collection
collection = resolve(bucketDir, bucket)
// Rule-less writes outside buckets fall back to the filer's default
// collection, so a bucket resolving there shares it with them.
// collection, so a bucket resolving there shares it with them. The
// system metadata-log collection (when explicitly redirected via
// filer.options.metaLog.collection) is in the same boat: it backs internal
// log volumes, so a bucket that resolves there must never drop it
// either.
if collection == f.metaLogCollection {
return ""
}
if f.metaLogTargetCollection != "" && collection == f.metaLogTargetCollection {
return ""
}
// A rule whose prefix escapes the bucket can route other paths into the
// same collection, including prefixes nested under surviving buckets.
+13 -2
View File
@@ -59,13 +59,24 @@ func (f *Filer) resolveMetadataLogAssignDiskType(targetFile string) (string, *fi
return util.Nvl(rule.DiskType, f.DefaultDiskType), rule
}
// metaLogCollectionFor resolves the system metadata log's collection: the
// filer.options.metaLog.collection override, then the filer default, then the
// matched storage rule.
func (f *Filer) metaLogCollectionFor(ruleCollection string) string {
return util.Nvl(f.metaLogTargetCollection, f.metaLogCollection, ruleCollection)
}
func (f *Filer) metaLogReplicationFor(ruleReplication string) string {
return util.Nvl(f.metaLogTargetReplication, f.metaLogReplication, ruleReplication)
}
func (f *Filer) assignAndUpload(targetFile string, data []byte) (*operation.AssignResult, *operation.UploadResult, error) {
// assign a volume location
diskType, rule := f.resolveMetadataLogAssignDiskType(targetFile)
assignRequest := &operation.VolumeAssignRequest{
Count: 1,
Collection: util.Nvl(f.metaLogCollection, rule.Collection),
Replication: util.Nvl(f.metaLogReplication, rule.Replication),
Collection: f.metaLogCollectionFor(rule.Collection),
Replication: f.metaLogReplicationFor(rule.Replication),
DiskType: diskType,
WritableVolumeCount: rule.VolumeGrowthCount,
ExpectedDataSize: uint64(len(data)),
+137
View File
@@ -0,0 +1,137 @@
package filer
import (
"context"
"testing"
"github.com/seaweedfs/seaweedfs/weed/pb"
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"github.com/seaweedfs/seaweedfs/weed/util"
)
// The metadata log's collection resolution must let operators redirect the
// internal /topics/.system/log chunks into their own collection without
// touching where user data goes, and must be strictly backward compatible
// when the override is unset.
func TestMetaLogCollectionResolution(t *testing.T) {
tests := []struct {
name string
target string // filer.options.metaLog.collection override
filerDefault string // -collection flag value
ruleCollection string // storage rule matched on the log path
want string
}{
{
name: "no override: filer default wins (today's behaviour)",
target: "",
filerDefault: "mydata",
ruleCollection: "",
want: "mydata",
},
{
name: "no override anywhere: still empty (default collection)",
target: "",
filerDefault: "",
ruleCollection: "",
want: "",
},
{
name: "override wins over filer default",
target: "filer-meta",
filerDefault: "mydata",
ruleCollection: "",
want: "filer-meta",
},
{
name: "override wins over matched rule",
target: "filer-meta",
filerDefault: "",
ruleCollection: "rulecol",
want: "filer-meta",
},
{
name: "no override but rule set: rule used (unchanged fallback chain)",
target: "",
filerDefault: "",
ruleCollection: "rulecol",
want: "rulecol",
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
// Drive the same setter a running filer would use post-construction.
f := &Filer{}
f.metaLogTargetCollection = tc.target
f.metaLogCollection = tc.filerDefault
if got := f.metaLogCollectionFor(tc.ruleCollection); got != tc.want {
t.Errorf("metaLogCollectionFor(%q) = %q, want %q", tc.ruleCollection, got, tc.want)
}
})
}
}
func TestMetaLogReplicationResolution(t *testing.T) {
f := &Filer{}
f.metaLogTargetReplication = "110"
f.metaLogReplication = "010"
if got := f.metaLogReplicationFor("001"); got != "110" {
t.Errorf("override must win, got %q", got)
}
f.metaLogTargetReplication = ""
if got := f.metaLogReplicationFor("001"); got != "010" {
t.Errorf("filer replication must win when no override, got %q", got)
}
f.metaLogReplication = ""
if got := f.metaLogReplicationFor("001"); got != "001" {
t.Errorf("rule replication must be used last, got %q", got)
}
}
// TestViperReadsMetaLogOverrides proves the exact viper keys the docs promise
// are read into the fields a running filer uses.
func TestViperReadsMetaLogOverrides(t *testing.T) {
v := util.GetViper()
v.Set("filer.options.metaLog.collection", "filer-meta")
v.Set("filer.options.metaLog.replication", "100")
defer func() {
// Reset so other tests see the unset default.
v.Set("filer.options.metaLog.collection", "")
v.Set("filer.options.metaLog.replication", "")
}()
f := NewFiler(pb.ServerDiscovery{}, nil, "", "", "", "", "", 255, nil)
if got := f.metaLogTargetCollection; got != "filer-meta" {
t.Fatalf("NewFiler did not read the collection override: %q", got)
}
if got := f.metaLogTargetReplication; got != "100" {
t.Fatalf("NewFiler did not read the replication override: %q", got)
}
}
// TestBucketCollectionKeepsMetaLogTargetCollection mirrors the guarantee that
// protects the filer's default collection: when a bucket resolves to the
// collection the system metadata log was redirected to, deleting the bucket
// must not drop that collection (it backs internal log volumes).
func TestBucketCollectionKeepsMetaLogTargetCollection(t *testing.T) {
f, store, master := newFilerWithFakeMaster(t)
// Operator redirected the meta log to its own collection...
f.metaLogTargetCollection = "filer-meta"
// ...and a bucket happens to resolve to that very collection.
f.FilerConf.SetLocationConf(&filer_pb.FilerConf_PathConf{
LocationPrefix: "/buckets/a",
Collection: "filer-meta",
})
seedBucket(t, store, util.FullPath("/buckets/a"))
if err := f.DeleteEntryMetaAndData(context.Background(), "/buckets/a", true, false, true, false, nil, 0); err != nil {
t.Fatalf("DeleteEntryMetaAndData: %v", err)
}
select {
case call := <-master.calls:
t.Fatalf("the meta-log target collection was deleted by a bucket delete: %q", call.name)
default:
}
}