Files
Nick Jamesandpc.yu 3841e6f299 fix(plugins): drain the dsh pending queue in-process so a transient write failure self-heals (#4779)
* fix(plugins): drain the pending queue in-process so a transient write failure self-heals

The dsh memory plugin latches capture and commit on the first retryable
write failure (hasPendingWrites) and only reset the latch at session
init, so the long-lived dsh process stayed stuck until restart.

Add a per-process single-flight drainer (default 60s, env
OPENVIKING_PENDING_DRAIN_INTERVAL_MS) that follows the session-start
flow: probe health, replay the queue without consuming retry budgets,
then re-derive every session's latch from the queue. replayPending gains
an optional consumeRetries flag (default true, byte-compatible):
drainers release a failed claim back to its original filename instead of
incrementing the retry count, so the session-start path keeps owning all
retry accounting and D4 deletions. Latch and health transitions are
logged once per flip for observability.

* test(plugins): cover the drainer and non-consuming replay mode

Add pending-queue coverage for consumeRetries:false (retryable failures
stay retryable and ordered, non-retryable and exhausted entries still
delete, commitSession failures keep the run going, default mode
unchanged) and runtime drainer coverage (recovery clears the latch,
outages keep it and leave entries retryable, empty queue means zero
HTTP, commit resumes after the drain, single-flight, per-session latch
isolation, interval wiring with env fallback).

---------

Co-authored-by: pc.yu <nick@fourieralpha.com>
2026-09-10 13:26:46 +08:00
..