Addressing PR #9382 review:
- Data race on lastIoError: guard lastIoError + lastIoErrorCount with a
RWMutex and expose them through note/clear/get helpers so the
heartbeat reader sees a consistent snapshot. Verified with -race.
- Collection-size accounting: when a volume is quarantined for sustained
EIO, skip the entire per-volume bookkeeping (`continue`) instead of
flipping shouldDeleteVolume — the old branch subtracted a size that
was never added, dragging the collection gauge to zero / negative.
- Recoverability: MarkVolumeWritable now also calls clearIoError so an
operator can rejoin a quarantined replica. The next failed op
re-arms the streak if the disk is still bad.
- Non-EIO streak break: a non-EIO error (e.g. ENOSPC) now resets the
consecutive-EIO counter, so a sequence EIO,EIO,ENOSPC,EIO is treated
as a streak of one — the counter only tracks consecutive EIOs.
Reads already call checkReadWriteError (volume_read.go), so successful
reads also clear the streak — no change needed there.