mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-10 00:07:44 +02:00
* pb: let a heartbeat carry only the volumes that changed A partial list cannot travel in volumes: a master that did not understand it would read the absences as deletions. So changes get their own field, used only once the master has said it compares digests and can tell when it has fallen behind. * master: apply the volumes a heartbeat reports as changed Only the named volumes are touched. A full report says the server holds exactly these; a changed report says nothing about the ones it leaves out, so absence must not read as removal. Also advertises that the master compares digests, which is what lets a server stop sending its whole list. Advertising it once per connection means a server reconnecting to a master that does not is back to full lists straight away. * volume: send only the volumes that changed once the master accepts them The whole list goes on every heartbeat until the master says it compares digests, and again whenever it asks, so a master that cannot tell when it has fallen behind never has to. has_no_volumes stays derived from a full list alone. Deriving it from what a heartbeat happens to carry would make a quiet one read as a server that had lost every volume, and the master would drop them all. The digest still covers every volume held rather than the ones sent, which is what lets the master confirm that applying the changes left it current. Reporting state is per-connection: a server that reconnects, or reaches a different master, starts again from the full list. * volume: let the zero reporting state stand for having told no master anything A Store built as a literal, which tests do, left the reporting state nil and panicked on the first heartbeat. As a value its zero form already means nothing has been reported to anyone, which is exactly the state that sends the whole list. * rust: send only the volumes that changed once the master accepts them Mirrors the Go volume server, with one hazard the Go side does not have: mount and unmount deltas here are derived by diffing successive heartbeats, so a heartbeat that carries a partial list would report every volume it left out as unmounted. Collecting now returns the full set alongside the message, and every site that diffs uses that rather than what went on the wire. * volume: do not let a full-list request be lost to the heartbeat it raced The request arrived while a heartbeat was already being built as a delta, and committing that heartbeat cleared it, so the master waited for another digest mismatch before asking again. Count the requests and clear only the one the heartbeat answered. * rust: stop marking volumes reported by a heartbeat that is thrown away The state-notify path collected a heartbeat only to diff its volume list, then sent a delta message of its own and dropped the one it had collected. Once collecting recorded what the master had been told, every mount or unmount silently marked the changed volumes as sent, and the master learned of them only after a digest mismatch. Snapshotting no longer records anything, and no longer expires ec volumes whose deletion that path was already discarding. * master: announce only the volumes a change actually brought Every changed volume was broadcast as a new location. Volumes grow constantly and growth moves no location, so on a busy cluster that told every connected client about volumes it could already reach, filling bounded broadcast queues and pushing out the topology updates that matter. * master: ask for the full list when only one can repair the master Delta heartbeats stop the full report, and with it the only thing that re-registers a volume the lookup index lost. The volume server cannot see that divergence and its digest cannot show it, so the master now checks its own two indexes agree and asks for the list when they do not. A node reporting one volume id twice is kept on full lists for the same reason rather than merely skipped: its digest can never be verified, so nothing else would tell the master what it had stopped holding. * master: keep the volume options on every heartbeat response A volume server takes them from whatever response arrives, and preallocate is a bare bool with no way to tell off from unmentioned. A response sent to ask for the volume list therefore turned preallocation off until the server reconnected. Responses sent mid-stream now start from the configured options rather than being built field by field. * master: announce a volume the lookup index had lost Repairing the index makes the volume servable again, but clients were told it went when the node dropped out and nothing told them otherwise: the disk map still held it, so it did not count as an arrival. Reaching the lookup index is what makes a volume servable, so recovering an entry there is an arrival as far as clients are concerned, on both the full report and the changed-volume path.
148 lines
6.0 KiB
Go
148 lines
6.0 KiB
Go
package weed_server
|
|
|
|
import (
|
|
"testing"
|
|
|
|
"github.com/seaweedfs/seaweedfs/weed/pb/master_pb"
|
|
"github.com/seaweedfs/seaweedfs/weed/sequence"
|
|
"github.com/seaweedfs/seaweedfs/weed/storage"
|
|
"github.com/seaweedfs/seaweedfs/weed/storage/needle"
|
|
"github.com/seaweedfs/seaweedfs/weed/storage/super_block"
|
|
"github.com/seaweedfs/seaweedfs/weed/storage/types"
|
|
"github.com/seaweedfs/seaweedfs/weed/topology"
|
|
)
|
|
|
|
func changedTestCluster(t *testing.T) (*topology.Topology, *topology.DataNode) {
|
|
t.Helper()
|
|
topo := topology.NewTopology("test", sequence.NewMemorySequencer(), 32*1024*1024*1024, 5, false)
|
|
dn := topo.GetOrCreateDataCenter("dc1").GetOrCreateRack("rack1").
|
|
GetOrCreateDataNode("127.0.0.1", 8080, 18080, "", "", map[string]uint32{"": 100})
|
|
return topo, dn
|
|
}
|
|
|
|
func changedTestVolume(id uint32, size uint64) *master_pb.VolumeInformationMessage {
|
|
return &master_pb.VolumeInformationMessage{
|
|
Id: id, Size: size, Collection: "c", Version: 3, FileCount: 1,
|
|
}
|
|
}
|
|
|
|
// Applying only the volumes a heartbeat named has to leave the master holding
|
|
// what the server holds, or the digest it is checked against means nothing.
|
|
func TestChangedVolumesBringTheMasterCurrent(t *testing.T) {
|
|
topo, dn := changedTestCluster(t)
|
|
topo.SyncDataNodeRegistration([]*master_pb.VolumeInformationMessage{
|
|
changedTestVolume(1, 1024), changedTestVolume(2, 1024), changedTestVolume(3, 1024),
|
|
}, dn)
|
|
|
|
grown := changedTestVolume(2, 8192)
|
|
topo.ApplyVolumeChanges([]*master_pb.VolumeInformationMessage{grown}, dn)
|
|
|
|
reference, referenceNode := changedTestCluster(t)
|
|
reference.SyncDataNodeRegistration([]*master_pb.VolumeInformationMessage{
|
|
changedTestVolume(1, 1024), grown, changedTestVolume(3, 1024),
|
|
}, referenceNode)
|
|
|
|
if dn.VolumeDigest() != referenceNode.VolumeDigest() {
|
|
t.Errorf("after applying the change the master digests %d, the server reports %d",
|
|
dn.VolumeDigest(), referenceNode.VolumeDigest())
|
|
}
|
|
}
|
|
|
|
// Silence about a volume in a changed-only heartbeat says nothing about whether
|
|
// the server still has it, unlike a full report.
|
|
func TestChangedVolumesDoNotRemoveUnmentionedVolumes(t *testing.T) {
|
|
topo, dn := changedTestCluster(t)
|
|
topo.SyncDataNodeRegistration([]*master_pb.VolumeInformationMessage{
|
|
changedTestVolume(1, 1024), changedTestVolume(2, 1024),
|
|
}, dn)
|
|
|
|
topo.ApplyVolumeChanges([]*master_pb.VolumeInformationMessage{changedTestVolume(1, 4096)}, dn)
|
|
|
|
if _, err := dn.GetVolumesById(needle.VolumeId(2)); err != nil {
|
|
t.Errorf("a volume the heartbeat did not mention was dropped: %v", err)
|
|
}
|
|
if !dn.HasConsistentVolumeIndex() {
|
|
t.Error("applying changes left the lookup index disagreeing with the disks")
|
|
}
|
|
}
|
|
|
|
// A volume the master has never seen can arrive as a change.
|
|
func TestChangedVolumesRegisterUnknownVolumes(t *testing.T) {
|
|
topo, dn := changedTestCluster(t)
|
|
topo.SyncDataNodeRegistration([]*master_pb.VolumeInformationMessage{changedTestVolume(1, 1024)}, dn)
|
|
|
|
topo.ApplyVolumeChanges([]*master_pb.VolumeInformationMessage{changedTestVolume(9, 1024)}, dn)
|
|
|
|
if locations := topo.Lookup("c", needle.VolumeId(9)); len(locations) != 1 {
|
|
t.Errorf("a volume first seen as a change is not servable: %v", locations)
|
|
}
|
|
if !dn.HasConsistentVolumeIndex() {
|
|
t.Error("applying changes left the lookup index disagreeing with the disks")
|
|
}
|
|
}
|
|
|
|
// Volumes grow constantly, and a growth moves no location. Telling every
|
|
// connected client about each one would flood bounded broadcast queues and push
|
|
// out the topology updates that do matter.
|
|
func TestChangedVolumesAnnounceOnlyArrivals(t *testing.T) {
|
|
topo, dn := changedTestCluster(t)
|
|
topo.SyncDataNodeRegistration([]*master_pb.VolumeInformationMessage{
|
|
changedTestVolume(1, 1024), changedTestVolume(2, 1024),
|
|
}, dn)
|
|
|
|
grown := topo.ApplyVolumeChanges([]*master_pb.VolumeInformationMessage{
|
|
changedTestVolume(1, 8192), changedTestVolume(2, 9216),
|
|
}, dn)
|
|
if len(grown) != 0 {
|
|
t.Errorf("volumes that only grew were announced as new locations: %v", grown)
|
|
}
|
|
|
|
arrived := topo.ApplyVolumeChanges([]*master_pb.VolumeInformationMessage{
|
|
changedTestVolume(1, 16384), changedTestVolume(7, 1024),
|
|
}, dn)
|
|
if len(arrived) != 1 || arrived[0].Id != needle.VolumeId(7) {
|
|
t.Errorf("expected only the volume that arrived, got %v", arrived)
|
|
}
|
|
}
|
|
|
|
// A node dropping out tells clients its volumes went. If the lookup index then
|
|
// loses one while the disk map keeps it, repairing the index has to announce it
|
|
// as well: nothing else will, and the clients that heard it go would never hear
|
|
// otherwise.
|
|
func TestRepairedLookupEntryIsAnnounced(t *testing.T) {
|
|
for _, tc := range []struct {
|
|
name string
|
|
apply func(*topology.Topology, *topology.DataNode, *master_pb.VolumeInformationMessage) []storage.VolumeInfo
|
|
}{
|
|
{"ViaChanges", func(topo *topology.Topology, dn *topology.DataNode, v *master_pb.VolumeInformationMessage) []storage.VolumeInfo {
|
|
return topo.ApplyVolumeChanges([]*master_pb.VolumeInformationMessage{v}, dn)
|
|
}},
|
|
{"ViaFullList", func(topo *topology.Topology, dn *topology.DataNode, v *master_pb.VolumeInformationMessage) []storage.VolumeInfo {
|
|
announced, _ := topo.SyncDataNodeRegistration([]*master_pb.VolumeInformationMessage{v}, dn)
|
|
return announced
|
|
}},
|
|
} {
|
|
t.Run(tc.name, func(t *testing.T) {
|
|
topo, dn := changedTestCluster(t)
|
|
volume := changedTestVolume(1, 1024)
|
|
topo.SyncDataNodeRegistration([]*master_pb.VolumeInformationMessage{volume}, dn)
|
|
|
|
rp, _ := super_block.NewReplicaPlacementFromString("000")
|
|
vl := topo.GetVolumeLayout("c", rp, needle.EMPTY_TTL, types.HardDriveType)
|
|
vl.SetVolumeUnavailable(dn, needle.VolumeId(1))
|
|
if locations := topo.Lookup("c", needle.VolumeId(1)); locations != nil {
|
|
t.Fatalf("expected the volume to be unservable, got %v", locations)
|
|
}
|
|
|
|
announced := tc.apply(topo, dn, changedTestVolume(1, 1024))
|
|
|
|
if locations := topo.Lookup("c", needle.VolumeId(1)); len(locations) != 1 {
|
|
t.Fatalf("the repair did not make the volume servable again: %v", locations)
|
|
}
|
|
if len(announced) != 1 || announced[0].Id != needle.VolumeId(1) {
|
|
t.Errorf("the repair was not announced to clients, so those told it went stay stale: %v", announced)
|
|
}
|
|
})
|
|
}
|
|
}
|