Files
seaweedfs/weed/storage/store_load_balancing_test.go
T
Chris Lu 3bd218e030 volume: cut idle memory at high volume counts (#10861)
* volume: start a volume's batch write worker on first use

Mounting a volume started a goroutine parked on a 128-slot channel, plus
the 128-entry batch slice it had already allocated. That is around 6.7KB
per volume the server pays whether or not the volume ever takes a write:
7231 bytes per mounted volume, of which 4101 is goroutine stack.

Only a write that asks for fsync ever reaches the worker, and a
remote-tiered or read-only volume never can. Create the channel and its
goroutine on the first such request instead, and let a write arriving
after Destroy fall back to the inline path rather than queue onto a
worker that has gone.

Measured over 20000 mounted volumes: 7231 -> 1269 bytes each.

* volume: update the heartbeat report state in place

Every heartbeat built a second map of what it was about to tell the
master, holding a freshly allocated short information message per volume,
then swapped it in over the old one -- and computed departures through a
third map of the live volume ids. A server holding 2M volumes rebuilt all
three every VolumePulsePeriod for a report that usually says nothing.

Number the heartbeats instead and mark the entry already held with the
pass that found the copy, so a quiet volume costs a map lookup and no
allocation. Departures are the entries a pass did not mark; the live-id
map is now built only when there are some, sized to them.

Measured over 10000 mounted volumes: 436 -> 196 bytes allocated per
volume per heartbeat.

* volume: fill one volume information message per heartbeat, not per volume

The heartbeat built a message for every volume held so it could hash it,
then dropped all but the few it had something to say about. At 2M volumes
that is 2M messages allocated every VolumePulsePeriod to send almost none
of them.

Fill a message the caller supplies instead, and replace it only when the
heartbeat keeps it, so a server with nothing to report fills the same one
all the way through.

Measured over 10000 mounted volumes: 196 -> 4 bytes allocated per volume
per heartbeat, and a heartbeat runs a third faster.

* volume: drop the per-volume trace from the heartbeat's status read

glog.V(4).Infof evaluates its arguments whether or not the verbosity is
on, so every volume boxed its id into a fresh interface slice on every
heartbeat: 759 of the 773 allocations a 1000-volume heartbeat made, for a
line that at this scale would print millions of unreadable rows.

Measured over 1000 mounted volumes: 4776 -> 1792 bytes and 759 -> 14
allocations per heartbeat, which no longer grows with the volume count.

* seaweed-volume: mirror the in-place heartbeat report state

Same change as the Go volume server: number the heartbeats and mark the
entry already held with the pass that found the copy, instead of building
a second map of hashes and swapping it in.

The volume snapshot must leave the reporting state as it found it, so it
keeps asking through changed() while a real heartbeat marks through
record().

* volume: refuse writes to a closed volume instead of dereferencing nil

Close and Destroy leave the needle map and data backend nil, but a caller
that already holds the volume can still reach the write path, where both
are used unguarded: a write racing a volume deletion took the server down.
syncDelete has always checked; syncWrite and the batch worker had not.

Reachable before this series and now also from the inline fallback a
durable write takes when the worker has gone.

* seaweed-volume: guard the report state with one mutex, as Go does

The full-list flag and the generation that answers it have to move
together. Split across separate atomics they cannot: a request landing
between begin's two reads returns full == false with the generation it
just raised, and one landing between commit's read and its clear is
marked answered by a heartbeat that carried no list. Either way the
resend is dropped.

Neither is reachable today -- every caller reaches this through the
store's RwLock, the flag setters under a read lock and the heartbeat
build under a write lock, so they cannot interleave. The type should not
depend on that being true two files away, and Go holds a single mutex
over exactly these fields.

* test: build the servers under test to match the harness's offset size

The mixed Go/Rust suites run both servers against one dataset, so both
have to agree on the offset width. They did not: the harness built Go
with no tags, 4-byte offsets, while the Rust crate defaults to its 5bytes
feature, and the Rust server then refused the .vif the Go server had just
written -- "bytes_offset mismatch: found 4, expected 5".

Build each side to match the offset size the test binary itself was
compiled with, so a plain `go test` and one with -tags 5BytesOffset both
get a matched pair.
2026-08-21 13:04:56 -07:00

261 lines
7.2 KiB
Go

package storage
import (
"os"
"path/filepath"
"strconv"
"testing"
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
"github.com/seaweedfs/seaweedfs/weed/stats"
"github.com/seaweedfs/seaweedfs/weed/storage/needle"
"github.com/seaweedfs/seaweedfs/weed/storage/super_block"
"github.com/seaweedfs/seaweedfs/weed/storage/types"
"github.com/seaweedfs/seaweedfs/weed/util"
)
// newTestStore creates a test store with the specified number of directories
func newTestStore(t testing.TB, numDirs int) *Store {
tempDir := t.TempDir()
var dirs []string
var maxCounts []int32
var minFreeSpaces []util.MinFreeSpace
var diskTypes []types.DiskType
for i := 0; i < numDirs; i++ {
dir := filepath.Join(tempDir, "dir"+strconv.Itoa(i))
os.MkdirAll(dir, 0755)
dirs = append(dirs, dir)
maxCounts = append(maxCounts, 100) // high limit
minFreeSpaces = append(minFreeSpaces, util.MinFreeSpace{})
diskTypes = append(diskTypes, types.HardDriveType)
}
diskIOProbeConfig := stats.DefaultDiskIOProbeConfig()
store := NewStore(nil, "localhost", 8080, 18080, "http://localhost:8080", "",
dirs, maxCounts, minFreeSpaces, "", NeedleMapInMemory, diskTypes, nil, 3, diskIOProbeConfig)
// Consume channel messages to prevent blocking
done := make(chan bool)
go func() {
for {
select {
case <-store.NewVolumesChan:
case <-done:
return
}
}
}()
t.Cleanup(func() {
store.Close()
close(done)
})
return store
}
func TestLocalVolumesLen(t *testing.T) {
testCases := []struct {
name string
totalVolumes int
remoteVolumes int
expectedLocalCount int
}{
{
name: "all local volumes",
totalVolumes: 5,
remoteVolumes: 0,
expectedLocalCount: 5,
},
{
name: "all remote volumes",
totalVolumes: 5,
remoteVolumes: 5,
expectedLocalCount: 0,
},
{
name: "mixed local and remote",
totalVolumes: 10,
remoteVolumes: 3,
expectedLocalCount: 7,
},
{
name: "no volumes",
totalVolumes: 0,
remoteVolumes: 0,
expectedLocalCount: 0,
},
}
for _, tc := range testCases {
t.Run(tc.name, func(t *testing.T) {
diskLocation := &DiskLocation{
volumes: make(map[needle.VolumeId]*Volume),
}
// Add volumes
for i := 0; i < tc.totalVolumes; i++ {
vol := &Volume{
Id: needle.VolumeId(i + 1),
volumeInfo: &volume_server_pb.VolumeInfo{},
}
// Mark some as remote
if i < tc.remoteVolumes {
vol.hasRemoteFile.Store(true)
vol.volumeInfo.Files = []*volume_server_pb.RemoteFile{
{BackendType: "s3", BackendId: "test", Key: "test-key"},
}
}
diskLocation.volumes[vol.Id] = vol
}
result := diskLocation.LocalVolumesLen()
if result != tc.expectedLocalCount {
t.Errorf("Expected LocalVolumesLen() = %d; got %d (total: %d, remote: %d)",
tc.expectedLocalCount, result, tc.totalVolumes, tc.remoteVolumes)
}
})
}
}
func TestVolumeLoadBalancing(t *testing.T) {
testCases := []struct {
name string
locations []locationSetup
expectedLocations []int // which location index should get each volume
}{
{
name: "even distribution across empty locations",
locations: []locationSetup{
{localVolumes: 0, remoteVolumes: 0},
{localVolumes: 0, remoteVolumes: 0},
{localVolumes: 0, remoteVolumes: 0},
},
expectedLocations: []int{0, 1, 2, 0, 1, 2}, // round-robin
},
{
name: "prefers location with fewer local volumes",
locations: []locationSetup{
{localVolumes: 5, remoteVolumes: 0},
{localVolumes: 2, remoteVolumes: 0},
{localVolumes: 8, remoteVolumes: 0},
},
expectedLocations: []int{1, 1, 1}, // all go to location 1 (has fewest)
},
{
name: "ignores remote volumes in count",
locations: []locationSetup{
{localVolumes: 2, remoteVolumes: 10}, // 2 local, 10 remote
{localVolumes: 5, remoteVolumes: 0}, // 5 local
{localVolumes: 3, remoteVolumes: 0}, // 3 local
},
// expectedLocations: []int{0, 0, 2}
// Explanation:
// 1. Initial local counts: [2, 5, 3]. First volume goes to location 0 (2 local, ignoring 10 remote).
// 2. New local counts: [3, 5, 3]. Second volume goes to location 0 (first with min count 3).
// 3. New local counts: [4, 5, 3]. Third volume goes to location 2 (3 local < 4 local).
expectedLocations: []int{0, 0, 2},
},
{
name: "balances when some locations have remote volumes",
locations: []locationSetup{
{localVolumes: 1, remoteVolumes: 5},
{localVolumes: 1, remoteVolumes: 0},
{localVolumes: 0, remoteVolumes: 3},
},
// expectedLocations: []int{2, 0, 1}
// Explanation:
// 1. Initial local counts: [1, 1, 0]. First volume goes to location 2 (0 local).
// 2. New local counts: [1, 1, 1]. Second volume goes to location 0 (first with min count 1).
// 3. New local counts: [2, 1, 1]. Third volume goes to location 1 (next with min count 1).
expectedLocations: []int{2, 0, 1},
},
}
for _, tc := range testCases {
t.Run(tc.name, func(t *testing.T) {
// Create test store with multiple directories
store := newTestStore(t, len(tc.locations))
// Pre-populate locations with volumes
for locIdx, setup := range tc.locations {
location := store.Locations[locIdx]
vidCounter := 1000 + locIdx*100 // unique volume IDs per location
// Add local volumes
for i := 0; i < setup.localVolumes; i++ {
vol := createTestVolume(needle.VolumeId(vidCounter), false)
location.SetVolume(vol.Id, vol)
vidCounter++
}
// Add remote volumes
for i := 0; i < setup.remoteVolumes; i++ {
vol := createTestVolume(needle.VolumeId(vidCounter), true)
location.SetVolume(vol.Id, vol)
vidCounter++
}
}
// Create volumes and verify they go to expected locations
for i, expectedLoc := range tc.expectedLocations {
volumeId := needle.VolumeId(i + 1)
err := store.AddVolume(volumeId, "", NeedleMapInMemory, "000", "",
0, needle.GetCurrentVersion(), 0, types.HardDriveType, 3)
if err != nil {
t.Fatalf("Failed to add volume %d: %v", volumeId, err)
}
// Find which location got the volume
actualLoc := -1
for locIdx, location := range store.Locations {
if _, found := location.FindVolume(volumeId); found {
actualLoc = locIdx
break
}
}
if actualLoc != expectedLoc {
t.Errorf("Volume %d: expected location %d, got location %d",
volumeId, expectedLoc, actualLoc)
// Debug info
for locIdx, loc := range store.Locations {
localCount := loc.LocalVolumesLen()
totalCount := loc.VolumesLen()
t.Logf(" Location %d: %d local, %d total", locIdx, localCount, totalCount)
}
}
}
})
}
}
// Helper types and functions
type locationSetup struct {
localVolumes int
remoteVolumes int
}
func createTestVolume(vid needle.VolumeId, isRemote bool) *Volume {
vol := &Volume{
Id: vid,
SuperBlock: super_block.SuperBlock{},
volumeInfo: &volume_server_pb.VolumeInfo{},
}
if isRemote {
vol.hasRemoteFile.Store(true)
vol.volumeInfo.Files = []*volume_server_pb.RemoteFile{
{BackendType: "s3", BackendId: "test", Key: "remote-key-" + strconv.Itoa(int(vid))},
}
}
return vol
}