mirror of
https://github.com/sandbaseai/sandbase-harness.git
synced 2026-09-28 14:13:23 +08:00
The claim is the first request of every iteration and was the only one still issued once with no bound and no error handling: ```ts const item = await claimWorkItem(config); ``` Two consequences followed from that line, and they are one behaviour seen from two sides - **the claim step could stop the worker**, either by parking it or by ending it. Unbounded, a claim the runtime never answered parked the process for ever. The stall was silent and total: no item, no report, no retry, and no message, because unlike a renewal there is no timer keeping it alive and so not even a growing pile to notice. Measured before the fix, the reproduction did not fail an assertion - it consumed the runner's full 30s timeout, which is what a permanent stall looks like from a test. Uncaught, a claim that *failed* was worse: the error escaped the poll loop and rejected the command, so a runtime that blinked killed every worker pointed at it. That one measured in 9ms. They are fixed together, because a bound alone would have been a regression: the timeout it raises would have taken the same uncaught path and turned the stall into a crash. The behaviour is that an item is run only when the claim produced one, and a claim that produced none - refused, failed, or never answered - is logged and the worker polls again after `--interval-ms`. `--claim-timeout-ms` (default 10000, minimum 1) bounds the request, and an expired bound is reported as the machine-readable `work_claim_unconfirmed` with the bound it waited, so "the runtime is not answering" is distinguishable from "the runtime is not there". A permanently wrong credential now reports at a steady rate rather than exiting, which is a deliberate change from a fast crash to a slow loop and is stated in the docs. The honest limit is documented where the decision is made rather than only in docs: a claim whose response was lost may still have created the row, so that item is stranded until its lease lapses. It is not lost, because its `accepted_at` is still null - it stays `queued` and the sweep re-hands it, which is the "unaccepted intent stays reclaimable" property doing its job. This also invalidated a test technique of my own from the delivery work: that loop's stopping condition was a sentinel thrown from the second claim, which is now caught and retried like any other claim error, so the test would poll for ever. It now stops the loop from the loop's own timer instead, which keeps the "reached a second claim" proof and additionally shows the loop got back to sleeping. Closes #661.
Documentation
This directory contains public, project-owned documentation for
managed-agents. It is written for users, contributors, and operators of the
open-source runtime.
Start Here
| Document | Audience | Contents |
|---|---|---|
| Installation | Users and operators | Install options, model configuration, startup flags, and health checks. |
| MiniMax | MiniMax users | Regional endpoints, model IDs, settings, and verification. |
| Usage Guide | Users and integrators | Workspace layout, Console workflows, sessions, sandbox backends, resources, credentials, memory, and SDK usage. |
| Showcase | New users and integrators | Three practical paths: auditable sessions, DeepSeek Harness integration, and controlled code execution. |
| API Reference | API and SDK integrators | HTTP endpoints, request shapes, response shapes, errors, and examples. |
| Versioned API Matrix | SDK authors and integrators | /v1 endpoint status, SDK coverage, CLI coverage, and compatibility gaps. |
| Skills | Agent builders | Skill package format, upload flow, validation rules, and agent references. |
| Requirements | Users and maintainers | Product scope, runtime guarantees, and release-facing requirements. |
| Technical Design | Contributors | Core concepts, data model, extension contracts, and API groups. |
| Architecture | Contributors and operators | System diagrams, data boundaries, session flow, and deployment modes. |
| Implementation Status | Contributors | Completed work, active work, planned items, and release checks. |
| Changelog | Users and maintainers | Release notes and first-release boundaries. |
Screenshots
Screenshots used by the README live in assets/:
Advanced / Optional
These documents are not part of the v1 quick-start path. Read them after the local SQLite + local filesystem runtime is working.
| Document | Audience | Contents |
|---|---|---|
| Deployment Examples | Operators | systemd, Docker Compose, running the runtime on Kubernetes, the RBAC for Kubernetes session sandboxes, self-hosted workers, and production checks. |
Release Gate
For a source checkout, the maintainer release gate is:
npm run release:check
It runs typecheck, tests, production builds, package dry-run, and CLI smoke
checks for managed-agents init plus examples/basic startup.
Documentation Rules
- Public docs describe this project only.
- Public docs avoid internal planning notes and external comparison material.
- Requirements describe observable behavior.
- Design docs describe stable architecture and extension points.
- Implementation status tracks current state without replacing issue tracking.