Files
Peter Steinberger c4954bf9c1 fix(state): prevent migration lease loss after database relocation (#159835)
* fix(state): prevent migration lease loss after database relocation

An inode-matched shared worker could retain the original admission of a
retired database path and acquire a migration lease in the wrong database.
Validate each reuse candidate separately from the current caller, join idle
actor retirement, and re-enter the ordinary caller admission path. Refuse
replacement while any callback still retains that actor.

The earlier e52420e1ee fix propagated an ended existing-schema candidate's
ordinary error. The follow-up 601bb5ba37 globally reclassified that error
and allowed ended callers to refresh their own admission. Keep schema policy,
error types, and caller checks unchanged; only candidate failure owns reuse
retirement. Restore the three real relocation cases and awaited migration
fixture cleanup. No schema, retention, permission, or update-driver changes.

Expanded proof also found two pre-existing test fixtures: supply the managed
Windows launcher sourcePath in Doctor service mocks, and await the catalog's
existing persisted-hydration signal before its first benchmark read. Preserve
all assertions, product deadlines, measurement loops, and performance limits.

Proof: restored cases fail 3/3 before the fix (90.57s wall); focused three-file
suites pass 74 tests (385.27s). The complete 464-file importer/glob inventory
passes 6,913 tests with 12 skips across local and Blacksmith runs. Original
44-file Linux CI replay passes 503 tests with two skips in 191.66s. Expanded
66-group proof covers 6,583 file executions and 96,306 cases; 65 groups are
green after the two fixture repairs. One Gateway recovery notification timeout
remains unexplained: its baseline 122-file group and instrumented candidate
208-test file pass. Record this limitation; do not claim a flake fix.
A local WAL observation timeout likewise did not reproduce in baseline and
candidate 16-file controls (394 passed, one skipped each), or its Linux group.

Single-worker test cost (Vitest / wrapper wall): shared-state 33 tests
240.84s / 257.60s; migration execution 23 tests 282.02s / 298.81s;
Doctor 72 tests 651.00s / 673.26s; catalog one test 132.86s / 140.34s.
Real-worker relocation and actor fencing require the real storage boundary;
fixture corrections add no cases. Core plus eight test graphs, focused
lint/format, docs, schema/ratchet/import-cycle guards, and diff checks pass.
Fresh independent Codex P2 review of all six files is scoped-clean.

* test(state): refresh relocation proof after task retirement

Main retired Tasks and TaskFlow while the stale-admission repair was under
validation. Exercise the same three real inode-relocation regressions through
pluginState.register with the production write-admission contract. Preserve
retained contents, retired-path absence, and active-callback refusal assertions.

The owner fix remains candidate-local: a current caller cannot renew a stale
actor's original admission. Unlike e52420e1ee, candidate ended-scope errors
cause idle retirement; unlike 601bb5ba37, schema-policy error types and ended
caller rejection remain unchanged. No production changes in this follow-up.

Forward the complete session-entry reader callback tuple in the media fixture,
expect the native orphan status after Tasks retirement, and reuse main's
maintenance error-cause assertion fix from 6ad6f1fbe8. Two lifecycle E2E cases now await the existing first-agent-call readiness
boundary after Tasks retirement removed the old mutation-tail drain. No deadline,
performance threshold, or substantive behavior assertion is relaxed.

Current admitted relocation control fails 3/3 against unpatched main (70.07s
wrapper), and passes 3/3 with the guard (26.28s). Single-worker file costs
(Vitest / wrapper): shared-state 32 cases 122.95s / 139.81s; media 12 cases
149.18s / 168.15s; orphan 27 cases 295.38s / 329.11s; maintenance 15 cases
70.70s / 117.61s. The existing three restored cases are the only added tests.

Refreshed required inventory plus media fixture: 396 files observed, 5,701
passed and 13 skipped cases after the documented repairs/controls; zero files
unexecuted. Initial/gap/Portal/lifecycle command walls total 2,777.81s. The final
four-file lifecycle replay passes 63/63 cases with original order and four
workers (558.76s including runtime rebuild). Its single-worker file passes
14/14 cases in 33.96s command / 11.675s native file time. Current containing
Linux CI shard: 59 files, 744 passed, two skipped, 181.92s command. Doctor/catalog
current-base follow-up: 73 passed, 146.14s command.

The refreshed inventory maps to 62 groups; only the containing current group
was replayed in full. Historical 66-group proof is retained separately; it is
not final-head coverage of every current group.

Nine type graphs passed (185.44s), with affected graph refreshes after final
fixture changes (8.93s, 13.10s and 10.19s). Lint, formatting, line guards and diff checks
pass. Independent Codex P2 review of the full ten-file candidate is scoped-clean
(283.45s). Historical WAL/Gateway timeouts and current post-build Portal timeout
remain unexplained after bounded controls; passing replays are not fixes.
2026-09-28 04:57:54 +00:00
..