Files
seaweedfs/weed/storage/store_duplicate_vid_test.go
T
10b0f2b8ad volume server: refuse the rest of a grouped run after a durable index failure (#11576)
* volume server: refuse the rest of a grouped run after a durable index failure

A durable write whose needle-map put fails stops the volume taking writes
(#10825): sent on its own, the next write then fails read only before it
appends. The grouped run from #11543 appends and syncs every entry before
publishing any, then kept publishing the entries after the failed one and
acked them once the shared .idx sync went through. When the failed put
tore its .idx row, the rows appended after it land off alignment, so the
next load parses them as garbage and the acked writes are gone.

Once a durable entry fails to publish, refuse every later entry of the run
with ReadOnly, as the per-needle path does. The entries before it stay
acked; their rows go down with the run's one .idx sync. The refused
records stay on the .dat unindexed, as the failed one does on its own.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: refuse a grouped entry staged as a cookie mismatch too

After a durable entry in a grouped run fails to index, the entries
after it are refused as they would be on their own. On its own an entry
meets check_writable before its cookie check, so one staged as a cookie
mismatch now gets the refusal too, instead of keeping its staging error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: trim a torn .idx row back so the next stays aligned

A failed write_index_entry can leave half a row in the .idx. With the
writer appending at the tail, every row written after it lands off
alignment and the next load parses them as garbage, so a write acked
behind a torn row does not come back. Trim the file back to
idx_file_offset on a failed append, in both needle maps, and cover it
with a test that writes past a torn row and reloads.

* volume server: refuse queued Go writes once a durable index update fails

processBatch kept writing after a failed nm.Put, and the single-write
path checked IsReadOnly only outside the volume lock. A durable write
whose index update fails now marks the volume noWriteOrDelete, and each
queued request is checked before it appends, so the ones after a failed
durable entry are refused the way a lone write is. Deletes get the same
noWriteOrDelete refusal a lone delete gets.

* volume server: refuse appends while a torn .idx row cannot be trimmed

When trimming back a half-written .idx row itself fails, the next append
would land after the torn bytes and every later row would parse off
alignment on load. Latch the map as torn and refuse appends until the
trim succeeds, on both CompactNeedleMap and RedbNeedleMap; the same
latch covers an orphan row that could not be trimmed after a failed
redb commit.

The .idx writer is now opened with write+append access so truncate_to
(set_len) works on Windows, where an append-only handle cannot trim.

* volume server: write .idx rows at idx_file_offset, not via append mode

Rust's OpenOptions on Windows strips FILE_WRITE_DATA whenever append is
set so the handle stays strictly append-only, which makes set_len fail -
the torn-row trim could never succeed there. Open the .idx writer with
plain write access and seek to idx_file_offset before each row, the same
positioned-write model the Go server uses.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 21:04:30 +08:00

131 lines
4.9 KiB
Go

package storage
import (
"testing"
"time"
"github.com/seaweedfs/seaweedfs/weed/storage/needle"
"github.com/seaweedfs/seaweedfs/weed/storage/super_block"
"github.com/seaweedfs/seaweedfs/weed/storage/types"
"github.com/stretchr/testify/require"
)
// A volume id can end up mounted on more than one disk of a server (a stale twin
// re-attached after a disk repair, since NewStore has no cross-disk duplicate
// guard). UnmountVolume must remove EVERY copy, not just the first match, or the
// stale twin survives and re-registers as the volume's content on the next
// heartbeat.
func TestUnmountVolumeRemovesAllDuplicateCopies(t *testing.T) {
store := newTestStore(t, 2)
const vid = needle.VolumeId(4242)
store.Locations[0].SetVolume(vid, createTestVolume(vid, false))
store.Locations[1].SetVolume(vid, createTestVolume(vid, false))
require.NoError(t, store.UnmountVolume(vid))
_, found0 := store.Locations[0].FindVolume(vid)
_, found1 := store.Locations[1].FindVolume(vid)
require.False(t, found0, "copy on disk 0 must be unmounted")
require.False(t, found1, "the stale twin on disk 1 must also be unmounted")
}
// DeleteVolume must likewise destroy every copy of a duplicate volume id, not just
// the first match, so the stale twin cannot survive the delete.
func TestDeleteVolumeRemovesAllDuplicateCopies(t *testing.T) {
store := newTestStore(t, 2)
const vid = needle.VolumeId(4243)
// Real volumes (not stubs) so Destroy can close and unlink them cleanly.
for _, loc := range store.Locations {
v, err := NewVolume(loc.Directory, loc.IdxDirectory, "", vid, NeedleMapInMemory,
&super_block.ReplicaPlacement{}, &needle.TTL{}, 0, needle.GetCurrentVersion(), 0, 0)
require.NoError(t, err)
loc.SetVolume(vid, v)
}
require.NoError(t, store.DeleteVolume(vid, false, false, false))
_, found0 := store.Locations[0].FindVolume(vid)
_, found1 := store.Locations[1].FindVolume(vid)
require.False(t, found0, "copy on disk 0 must be deleted")
require.False(t, found1, "the stale twin on disk 1 must also be deleted")
}
// A guarded delete must check every copy before destroying any: an earlier
// garbage copy must survive when a later duplicate still holds live data.
func TestDeleteVolumeGuardChecksAllCopiesBeforeDeleting(t *testing.T) {
store := newTestStore(t, 2)
const vid = needle.VolumeId(4244)
for _, loc := range store.Locations {
v, err := NewVolume(loc.Directory, loc.IdxDirectory, "", vid, NeedleMapInMemory,
&super_block.ReplicaPlacement{}, &needle.TTL{}, 0, needle.GetCurrentVersion(), 0, 0)
require.NoError(t, err)
loc.SetVolume(vid, v)
}
// Disk 0's copy is fully deleted; disk 1's twin still holds a live needle.
copy0, _ := store.Locations[0].FindVolume(vid)
copy1, _ := store.Locations[1].FindVolume(vid)
n := &needle.Needle{Id: types.Uint64ToNeedleId(1), Data: []byte("x")}
_, _, _, err := copy0.writeNeedle2(n, false, false, false)
require.NoError(t, err)
_, _, _, err = copy1.writeNeedle2(n, false, false, false)
require.NoError(t, err)
_, err = copy0.deleteNeedle2(n)
require.NoError(t, err)
err = store.DeleteVolume(vid, false, true, false)
require.ErrorIs(t, err, ErrVolumeNotEmpty)
_, found0 := store.Locations[0].FindVolume(vid)
_, found1 := store.Locations[1].FindVolume(vid)
require.True(t, found0, "the garbage copy must survive a refused delete")
require.True(t, found1, "the live copy must survive a refused delete")
}
// A guarded delete must serialize with writes in flight on any copy: while a
// copy's data lock is held (a write in progress), the delete cannot start
// destroying other copies, or a write landing between validation and removal
// would refuse the later copy and leave a partial delete.
func TestDeleteVolumeGuardWaitsForInFlightCopyWrite(t *testing.T) {
store := newTestStore(t, 2)
const vid = needle.VolumeId(4245)
for _, loc := range store.Locations {
v, err := NewVolume(loc.Directory, loc.IdxDirectory, "", vid, NeedleMapInMemory,
&super_block.ReplicaPlacement{}, &needle.TTL{}, 0, needle.GetCurrentVersion(), 0, 0)
require.NoError(t, err)
loc.SetVolume(vid, v)
}
copy1, _ := store.Locations[1].FindVolume(vid)
copy1.dataFileAccessLock.Lock()
done := make(chan error, 1)
go func() {
done <- store.DeleteVolume(vid, false, true, false)
}()
select {
case err := <-done:
t.Fatalf("guarded delete proceeded while a copy was locked for write: %v", err)
case <-time.After(200 * time.Millisecond):
}
// The write wins; the delete must see the new needle and refuse, leaving
// every copy intact.
n := &needle.Needle{Id: types.Uint64ToNeedleId(1), Data: []byte("x")}
_, _, _, err := copy1.doWriteRequest(n, false, false)
require.NoError(t, err)
copy1.dataFileAccessLock.Unlock()
require.ErrorIs(t, <-done, ErrVolumeNotEmpty)
_, found0 := store.Locations[0].FindVolume(vid)
_, found1 := store.Locations[1].FindVolume(vid)
require.True(t, found0, "the earlier copy must survive a refused delete")
require.True(t, found1, "the written copy must survive a refused delete")
}