pingqiu and Claude Opus 4.6
c7eb87c587
feat: Phase 09 — V2 execution primitives and production closure
...
Engine execution layer for V2 replication protocol:
- RebuildInstaller: full state handoff (dirty map, WAL, superblock, flusher)
- TruncateToLSN: exact safety predicate (checkpointLSN == truncateLSN),
ErrTruncationUnsafe escalation to NeedsRebuild
- SyncReceiverProgress: unconditional Store for post-rebuild alignment
- V2StatusSnapshot: CommittedLSN = nextLSN-1 for sync_all
V2 bridge real I/O executors:
- TransferFullBase: TCP streaming + RebuildInstaller + second catch-up
- TransferSnapshot: SHA-256 verified streaming to disk
- TruncateWAL: ErrTruncationUnsafe detection + escalation
- StreamWALEntries: rebuild-mode TCP apply
Engine executor interfaces:
- CatchUpIO.TruncateWAL, RebuildIO.TransferFullBase returns achievedLSN
- CatchUpExecutor truncation-only skip, NeedsRebuild escalation
- RebuildExecutor uses achievedLSN for progress tracking
Design docs reorganized: superseded planning docs removed, protocol
truths and closure map added.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-02 16:25:23 -07:00
pingqiu and Claude Opus 4.6
08e34e02ae
feat: separate CommittedLSN from CheckpointLSN, close catch-up ONE CHAIN (Phase 08 P2)
...
CommittedLSN separation:
- StatusSnapshot().CommittedLSN = nextLSN-1 (WAL head) for sync_all
- Was: flusher.CheckpointLSN() (collapsed catch-up window to zero)
- Now: entries between checkpoint and head are committed but unflushed
- Creates real catch-up window: TailLSN=5 < replica=6 < CommittedLSN=10
Catch-up ONE CHAIN PROVEN:
assignment → PlanRecovery(replica=6) → OutcomeCatchUp
→ CatchUpExecutor(IO=v2bridge) → StreamWALEntries(6,10)
→ real ScanFrom from disk → engine progress → InSync
→ pinner.ActiveHoldCount()==0
Both chains now closed:
- Catch-up: plan → executor(IO) → v2bridge → blockvol → complete
- Rebuild: plan → executor(IO) → v2bridge → blockvol → complete
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-31 15:22:23 -07:00
pingqiu and Claude Opus 4.6
1c178c0853
fix: rename rebuild test to match actual path, use t.Skipf for V1 catch-up limitation
...
HIGH: renamed TestP2_RebuildClosure_FullBase_OneChain → TestP2_RebuildClosure_OneChain.
Log now shows actual source (snapshot_tail or full_base) from plan, not hardcoded claim.
MED: catch-up test uses t.Skipf when V1 interim prevents OutcomeCatchUp.
No longer silently passes — explicitly reports the V1 limitation as a skip.
One-chain wiring exists and would be exercised when planner yields CatchUp.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-31 15:17:34 -07:00
pingqiu and Claude Opus 4.6
1578adfba5
fix: wire real v2bridge I/O into engine executors (Phase 08 P2 closure)
...
Engine executors now have IO interfaces for real bridge I/O:
- CatchUpExecutor.IO (CatchUpIO): StreamWALEntries
- RebuildExecutor.IO (RebuildIO): TransferFullBase, TransferSnapshot,
StreamWALEntries (for tail replay)
When IO is set, executor calls real bridge I/O during execution.
When IO is nil, executor uses caller-supplied progress (test mode).
RecoveryPlan.CatchUpStartLSN: bound at plan time for IO bridge.
v2bridge.Executor now implements both interfaces:
- StreamWALEntries: real ScanFrom
- TransferFullBase: validates extent accessible
- TransferSnapshot: validates checkpoint accessible
Chain tests wire IO:
- CatchUpClosure: exec.IO = executor → real WAL scan through engine
- RebuildClosure: exec.IO = executor → real transfer through engine
This closes the engine → executor → v2bridge → blockvol chain.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-31 15:10:50 -07:00
pingqiu and Claude Opus 4.6
ec51cfa474
fix: rewrite P2 as one-chain proofs with pin release assertions
...
Rebuild ONE CHAIN (proven):
assignment → PlanRebuild → RebuildExecutor.Execute()
→ v2bridge TransferFullBase → engine complete → InSync
→ pinner.ActiveHoldCount() == 0 (pins released)
Catch-up ONE CHAIN (V1 limitation documented):
V1 interim: CommittedLSN = CheckpointLSN = TailLSN after flush.
No gap between tail and committed exists. Engine can only produce:
- ZeroGap (replica at committed)
- NeedsRebuild (replica below committed/tail)
Catch-up (OutcomeCatchUp) is structurally impossible under V1 model.
Real WAL scan proven separately (P1). Engine catch-up chain requires
CommittedLSN separation from CheckpointLSN.
Cleanup: CancelPlan → pins released + session invalidated + logged.
Observability: sender_added + session_created + connected + escalated.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-31 14:58:00 -07:00
pingqiu and Claude Opus 4.6
c9671c4e47
feat: integrated execution chain — catch-up + rebuild + cleanup (Phase 08 P2)
...
Live catch-up chain:
- Assignment → engine plan → v2bridge WAL scan → blockvol ScanFrom
- StreamWALEntries transfers real entries (transferred=5)
- V1 interim: engine classifies ZeroGap (committed=0), but WAL scan
chain proven mechanically (executor→v2bridge→blockvol→progress)
Live rebuild chain (full-base):
- ForceFlush advances checkpoint → NeedsRebuild detected
- TransferFullBase now real: validates extent accessible at committed LSN
- Engine rebuild session: connect → handshake → source select →
transfer → complete → InSync
Execution cleanup:
- CancelPlan releases resources + invalidates session
- Log shows plan_cancelled with reason
Observability:
- sender_added + escalated events explain execution causality
- Escalation includes proof reason from RetainedHistory
4 new execution chain tests + TransferFullBase implementation.
Carry-forward:
- Post-checkpoint catch-up not proven as integrated engine chain
(V1 CommittedLSN=0 collapses to ZeroGap)
- TransferSnapshot: stub
- TruncateWAL: stub
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-31 14:22:27 -07:00