A claim is a lease on running a work item, not ownership of it: `claim()` hands
out an item with a 60s window and the row becomes claimable again the instant it
passes. The gap between claiming and starting is unbounded - the worker returns
from the claim, resolves its workdir, and may be a process that was paused or
descheduled - so an item whose window passed could be reclaimed and run by a
second worker while the first one was also running it. The tool call then
happened twice on the operator's machine.
A heartbeat does not close this. A renewal from a superseded holder is refused
as `not_claimed_by_worker`, which the CLI deliberately treats as a suspicion and
keeps running: right for an item already in flight, whose effect cannot be
un-run, and useless as a guard against starting one.
`POST /v1/x/worker/accept` is that guard. It is fenced on the item still being
claimed by the calling worker, unstopped, and inside its lease window, and only
a success authorizes execution. The lease fence is deliberately stricter than
`heartbeat`'s: a lapsed claim is reclaimable at any instant, so accepting one
would authorize a start on work the queue may hand to another worker in the same
instant, while a renewal cannot cause a second execution. Both a stopped item
and a lapsed lease answer `409` with the engine-neutral `work_lease_lost`,
because the instruction to the worker is the same - do not run it - and the
message says which of the two it was. A foreign holder gets `409` without a
code, since the work is alive for its owner, and an unknown id gets `404`.
A refusal is never a completion: `worker poll` does not run the item and does
not report a result, which would assert an effect that never happened.
Acceptance is recorded in a new `accepted_at` column (migration 046), cleared
when another worker reclaims the item and idempotent for the holder. Existing
rows stay NULL, reading as "held but never started" - the honest back-fill,
since there was no accept step for an earlier-build row to have been accepted
by, and inventing a timestamp would let the queue treat a worker that died
before starting as one that may already have run the item.
`WorkQueue.await` ended every failure with "work item <id> timed out", so a
caller could not tell apart three situations that need opposite responses: an
item nobody ever claimed, which is safe to submit again; an item an executor
held and never reported, where the effect may already have happened on the
operator's machine; and an item the session stopped, which is not wanted at all.
Through the session-error path an unrecognized code reported `unknown`, which
that documentation itself calls an invitation to retry the failures that must
not be retried.
The failure now carries a code, classified from the row the wait loop already
read. A stop is reported ahead of the item's state, because a stopped item that
nobody had claimed is still `pending`, and reporting that as "submit it again"
would send work back to a session that has ended. `pending` is the one case that
is provably replayable, since a row becomes `claimed` when it is handed out and
nothing moves it back. Everything else is the unknown outcome, a row that cannot
be read included: absence proves nothing about whether the work ran, and of the
two ways to be wrong, replaying an effect that already happened is the expensive
one.
`retry_status` classifies the two new codes rather than leaving them to fall
through, and a result reported before the deadline still resolves the wait,
stopped work included - the marker says the result is not wanted, not that the
effect did not happen.
`queue.stop()` is reached only through the sandbox release. A terminal session
is allowed to skip that release - `cleanup_pending` does, to retain the
workspace until child-tree cleanup can be proven - and a session that simply
finishes its turn never releases one at all. Work left unmarked in either case
was claimable, so a worker executed a tool call for a session that was over.
The claim predicate now asks the session directly. That also covers rows
already sitting unmarked in an existing database, where the insert-time check
cannot reach. The refusal is part of the selection as well as the guarded
update: refusing a row after selecting it would leave the queue holding work it
cannot hand over, starving workers of the items behind it.
The holder may still report a result that happened - the session ending says
the work is not wanted, not that the effect did not occur. Work for an unknown
session id, and for a status that is merely not running, stays claimable.
`WorkQueue.stop(sessionId)` marks the session's unfinished work, and
`enqueue()` then inserted new rows with `stopped_at` NULL unconditionally.
The stop is a decision about the session, so work that arrived afterwards - a
tool call from a turn that was still in flight, or one enqueued by the second
process this protocol exists to tolerate - was claimable, and a worker
executed it for a session that was already over.
The session's status is now read inside the same statement as the insert, so
there is no window between the check and the write, and the terminal statuses
are derived from the session state machine rather than listed again. An id
with no session row stays claimable: an unknown id is a caller's mistake
rather than a stop. A status that is merely not running stays claimable too -
paused, requires_action, and failed are all resumable, and `failed` has a
transition back to `running`.
Rows that were already queued remain the queue stop's business. The one path
where that leaves a gap - a terminal session that skips its sandbox release,
so nothing calls `queue.stop()` for it - is filed separately as #631.
* test(pi-launcher): wait for a complete result document, not for the path
The helper returned on existence, so a loaded run could parse a partial file.
* docs(changelog): record the fixture wait fix
Test-only change, recorded like the other test-only entries in the section.
* feat(sandbox): let a worker renew its claim lease
Only the holder can renew, so a long item is not reclaimed mid-flight.
* docs(contracts): register the heartbeat route in the route inventories
A mounted route is enumerated in three places; the gate named all three.
* docs(contracts): register the heartbeat route in the route inventories
A mounted route is enumerated in three places; the gate named all three.
* test(sandbox): pin that an unknown work item is 404, not 409
The route has two refusals and only the cross-worker one was driven.
* test(sandbox): pin that an unknown work item is 404, not 409
The route has two refusals and only the cross-worker one was driven. The error-code coverage fixture records the new not_found mention.
* test(pi): pin that suspending the heartbeat keeps the lease
The launcher suspends instead of releasing while a cleanup is pending, so the exclusion has to survive.
* test(pi): report a withdrawn lease before reading it
The existence check runs first so the probe names the missing lease instead of an ENOENT.
Every existing signature assertion is self-consistent: the verifier calls the signer, and the secret test compares the signer with itself. The signed content and the decoded key are claims about a receiver.
Presence is the weakest form of the retention claim. The delete records a termination and then the deletion itself, and the first expectation written for this failed because only one of the two writers was read.
The forward-only scan, the session-bound cursor, and the refusal of a cursor naming an absent event had no coverage; the last one would otherwise answer a request to continue with the whole log.
Two copies of the harness is where they start to disagree, and a difference in the harness reads as a difference in the runtime. The driver contract is asserted rather than described.
The unique-violation branch has no pre-check by design, and nothing checked that the refused write leaves the first record intact or that the constraint is scoped to the store.
pi_policy_mismatch and outcome_rubric_file_not_found are asserted only through their imported constants, so both published strings were unpinned: drift the value and every existing assertion still passes.
The code, the empty path, the warnings array on the failure path and the contrast with a schema-invalid body; the sibling route answers the same fault in a different shape, recorded rather than resolved.
The status was asserted; the envelope the fallback comment says clients depend on was not, nor the documented refusal to reflect the requested path. The non-reflection case reads the raw body so it cannot fail for a content-type reason.
The path decides which beta a request needs, with one documented exception; the code says what the regex matches, and only a request says whether the exception is reached.
Each covers sub-cases with opposite dispositions, so a stated value would be wrong for one of them; the test makes a symmetry-driven entry fail rather than land silently.
The declined-dialog behaviour was covered through the error class, but the published code a client switches on was asserted nowhere, so it could be renamed with every test still green.
A timeout and an unknown outcome both carry outcomeUnknown, and the transport states the command must not be retried blindly. Unknown invited that retry.
The code was absent from the retry-classification table, so a permanent frame-trust failure was published as unknown, which tells a client it might succeed on retry. The transport requires the opposite.
The unguarded-execution behaviour was covered through the error class, but the published code a client switches on was asserted nowhere, so it could be renamed with every test still green.
The code and the marker that survives serialization were asserted nowhere, so losing the marker would silently stop an outcome-unknown failure from being recognized across a boundary.
The frame-trust failure had no literal assertion, so drift that merged it with a retryable code would have told clients to resend a frame the runtime always refuses.
* test(api): pin pi_rpc_closed and its distinctness from pi_rpc_session_closed
The transport-closed code had no assertion, so the documented pair a caller branches on for retry disposition could collapse into one code unnoticed.
* test(api): assert the rpc closure pair against the constants, not two literals
The first version compared two literals, so collapsing the pair in rpc-wire.ts could not disconfirm it: the probe passed unchanged. Reviewed after the probe failed to fail, which also required regenerating the measured coverage fixture.
The code was asserted only through its constant, so the literal callers switch on could change unnoticed. The new test asserts the spelling, the 401/403 boundary, that 404 stays distinct, and that a status is never guessed from a message.
The model configuration failure carried an unasserted code, and its membership in the resumable model-failure set was untested, so a repairable configuration mistake could silently start terminating sessions.
The memory store-unavailable refusal was asserted by no test, so neither its wire code nor the store identifier in its message was pinned. The new unit test covers both and requires that two stores do not report the same message.