Files
Yao 2c3b0c808c fix(worker): confirm a claim before running the item, not only when taking it (#637)
A claim is a lease on running a work item, not ownership of it: `claim()` hands
out an item with a 60s window and the row becomes claimable again the instant it
passes. The gap between claiming and starting is unbounded - the worker returns
from the claim, resolves its workdir, and may be a process that was paused or
descheduled - so an item whose window passed could be reclaimed and run by a
second worker while the first one was also running it. The tool call then
happened twice on the operator's machine.

A heartbeat does not close this. A renewal from a superseded holder is refused
as `not_claimed_by_worker`, which the CLI deliberately treats as a suspicion and
keeps running: right for an item already in flight, whose effect cannot be
un-run, and useless as a guard against starting one.

`POST /v1/x/worker/accept` is that guard. It is fenced on the item still being
claimed by the calling worker, unstopped, and inside its lease window, and only
a success authorizes execution. The lease fence is deliberately stricter than
`heartbeat`'s: a lapsed claim is reclaimable at any instant, so accepting one
would authorize a start on work the queue may hand to another worker in the same
instant, while a renewal cannot cause a second execution. Both a stopped item
and a lapsed lease answer `409` with the engine-neutral `work_lease_lost`,
because the instruction to the worker is the same - do not run it - and the
message says which of the two it was. A foreign holder gets `409` without a
code, since the work is alive for its owner, and an unknown id gets `404`.

A refusal is never a completion: `worker poll` does not run the item and does
not report a result, which would assert an effect that never happened.

Acceptance is recorded in a new `accepted_at` column (migration 046), cleared
when another worker reclaims the item and idempotent for the holder. Existing
rows stay NULL, reading as "held but never started" - the honest back-fill,
since there was no accept step for an earlier-build row to have been accepted
by, and inventing a timestamp would let the queue treat a worker that died
before starting as one that may already have run the item.
2026-09-27 21:23:53 +08:00
..