A file or repository resource is materialized at a canonical absolute
in-sandbox path (`/mnt/session/uploads/...`, `/workspace/<repo>`). A backend
that refuses those paths could never serve such a session, but the session was
accepted with a 201 and failed later at provisioning, where the caller had no
way to connect the path error to the environment they chose.
Admission now decides: a session whose Environment resolves to a backend that
cannot serve the canonical roots is refused with 400, `invalid_request_error`,
and the new code `resource_not_mountable`, before anything is stored — no
session row, no resource instance, no event. The refusal is raised through
`SessionManager.create`, so it covers `POST /v1/sessions` and `POST /v1/runs`
in every response mode including `sse` (which answers the JSON error rather
than opening a stream), and through `SessionManager.assertResourcesMountable`
on `POST /v1/sessions/{id}/resources` for attaching a resource to an existing
session. An Environment that stops resolving on that route answers with its own
400 and code instead of escaping as a 500.
The refusal list is explicit — docker, kubernetes, self_hosted — in
`src/core/resources/resource-mountability.ts`, which owns the code, the mounted
resource types, the per-backend reason, and the guard the three routes map
through. Docker refuses every absolute path; kubernetes refuses the upload root
and its acceptance of the repository root was never exercised against a
cluster; a self-hosted worker maps an absolute path into its own root, so the
canonical root is not the runtime's to promise. A backend this build does not
ship keeps its previous behaviour, because refusing on a guess would reject a
provider that works.
With the gap closed at creation, `file-resources` and
`github-repository-materialization` are `supported` for the `local` backend
that serves them, and the canaries and guard blockers that recorded the gap are
removed together with the status they justified — the provider path checks they
rested on are kept as ordinary cases. Matrix reasons, the capability status
blocks, both contract documents, `docs/api.md`, `docs/api-matrix.md`,
`docs/usage.md`, the Console API reference, and `CHANGELOG.md` are updated in
the same change, and the error-code fixtures register the new code.
The append route decides from the Environment the session names, the same
authority creation uses, so the decision follows the Environment in both
directions rather than the sandbox already bound; that is documented rather
than solved.
Verification: typecheck, build, and package check pass; the suite passes with
277 files and 2553 tests (2 files and 17 tests skipped, the kubernetes suite
among them for lack of a reachable cluster). The new integration suite was
mutation-probed — with the four call sites stashed, 7 of its 9 cases fail — so
the assertions are load-bearing. No credential is echoed on a refusal. Docker
is exercised as a recording provider and no live model provider is involved,
because the behaviour under test is the admission decision.
A session's file and repository resources are written to paths the runtime
publishes, but nothing ever told the agent what those paths were: a mount it
was never told about is a mount it will not read.
A session that declares a `file` or `github_repository` resource now carries a
`# Session Resources` section naming each one -- a file by the sandbox path its
bytes were written to, derived by the same function the provisioning pass uses,
and a repository by URL, checkout, and mount path. On the `local` backend each
entry also carries the sandbox-relative spelling a shell needs, because a local
command resolves an absolute path against the host filesystem; a container
backend's own root is where the canonical path points, so the extra spelling is
added there and only there. The backend is read from the sandbox bound at
provision time rather than re-read from the Environment, so editing an
Environment after a session is bound cannot change what its instructions say.
Nothing here reads a credential: only the fields the contract publishes are
rendered, userinfo is stripped from a URL before printing, and an entry whose
shape is not understood -- an unrecognised checkout, a value carrying a newline
-- is dropped rather than guessed at or rewritten.
The two capability entries stay `partial`, since the container backends still
refuse the canonical roots and a session on one is still accepted before it
fails at provisioning, but their reason no longer claims the mount is
unannounced.
* feat(sandbox): resolve the canonical in-sandbox roots on the local backend
The runtime publishes absolute paths to an agent and to the host: uploaded file resources at /mnt/session/uploads/..., a mounted repository at /workspace/<repo>, spilled tool output at /mnt/session/tool_outputs/..., and session deliverables collected from /mnt/session/outputs/.... LocalSandboxProvider confined every file path to its own per-session directory, so on the local backend a session that declared a file or repository resource was accepted and then failed at provisioning, and a deliverable written under the canonical output root was never collected.
Both roots are now mapped into the sandbox directory by dropping the leading separator, and the mapped path goes through the same resolution, containment, and realpath checks as a relative one. The roots stay distinct directories rather than both aliasing the sandbox root: one shared root would put a repository named 'uploads' or 'outputs' into the directory the runtime reads uploads from, publishes deliverables from, and spills oversized tool output into.
* docs(sandbox): state the canonical roots' shell limit and split the container claims
An independent blind review of the path-mapping change found two claims that read as more than the code does: a command the agent runs cannot name a canonical absolute path (the mapping is applied to the file API and to an execute working directory, not to a command string, which runs as a host subprocess), and kubernetes does not refuse the canonical repository root (/workspace/<repo> resolves inside its own /workspace).
The provider header, the SandboxInstance documentation, the files and github-repository contracts, docs/api.md, docs/usage.md, and the CHANGELOG entry now say what each backend does per root, and two tests pin the behaviour: a command reads the file by its sandbox-relative spelling, and resolveWorkspacePath still accepts /workspace/<repo> so a future refusal fails the canary instead of leaving a wrong sentence in the contract. The pre-existing dangling-symlink write case is recorded in the provider header rather than claimed as covered.
`ManagedAgentsApiError` exposed only `status` and a message built from
`error.message`, so the runtime's error taxonomy was unreachable from the SDK:
API error 400: agent must be a standard agent id
The published envelope is structural, and the runtime builds it in one place
(`src/api/routes/sessions.ts:766`):
{"error":{"type":"invalid_request_error","code":"invalid_agent_ref","message":"..."}}
A caller is meant to branch on that identity. Without it, distinguishing a wrong
agent reference from a malformed body means substring-matching English prose that
is not a stable interface, and any rewording of a message is a silent breaking
change for the consumer that was forced to match on it. This closes the client
half of the error taxonomy (#464) and does not change the canonical type (D11).
* `error.type` and `error.code` carry the runtime's own values, and are
`undefined` when the response carried no envelope or named no specific cause
(`not_found` is a type, not a code under `invalid_request_error`).
* The message is deliberately **unchanged** — `API error 400: agent must be a
standard agent id` — so this is additive for anyone already matching on it.
Measured: only one assertion in the suite matches an error message
(`tests/integration/sdk.test.ts:251`, `/API error 404/`).
* The error body is now read exactly once. The previous implementation called
`res.json()` and then `res.text()` on the same `Response`, whose body is
single-use, so the non-JSON fallback could never succeed: a non-JSON failure
was reported as `statusText`, not as the body the runtime sent. The fallback
was dead code that looked alive.
* `requestText` — which backs artifact and file content plus metrics — goes
through the same reader. It used the raw body as the message, so a JSON
envelope was reported as raw JSON text with no fields readable.
* A JSON body that is not an envelope keeps the previous `statusText` fallback,
so only the envelope's fields are newly reachable; an unrelated JSON payload
cannot leak into a log line.
Verification: `npm run release:check` — see PR body for counts. The integration
test drives a real listener over real SQLite, so the codes are produced by the
routes: `invalid_agent_ref` and `agent_required` from `POST /v1/sessions`,
`not_found` from `GET /v1/sessions/{id}` and from the `requestText` path, and the
message pinned byte-for-byte so adding fields cannot alter the sentence existing
callers log. The non-JSON and malformed-envelope branches are covered by a
stubbed fetch, because no route in the suite answers a failure with a non-JSON
body.
Not covered: no Docker daemon, no Kubernetes cluster, and no model provider
credentials exist on this host, so no provider boundary is exercised;
`invalid_initial_events` was not reachable, because that route resolves a default
environment before it validates `initial_events` and answers
`Environment not found: env_default` with no code.
Reverse probes, each differing from the fix in one respect:
A. type/code dropped from the constructor -> see PR body
B. body read twice, as before -> see PR body
C. envelope required to be read from the
raw text rather than parsed -> see PR body
`docs/api-matrix.md:86` documents the `settings` group as covered — "Get, set
model boundary, and validate canonical runtime settings" — and
`docs/spec/tasks.md:144` ticks its three commands. They are implemented in
`src/cli/runtime-management-commands.ts`, the module the previous change
registered only halfway: these are the commands it deliberately left
unregistered rather than register broken (#466).
Driven against a live listener before the change:
settings get TypeError: Cannot read properties of undefined (reading 'metadata')
settings validate TypeError: result.checks is not iterable
settings set-model ManagedAgentsApiError: API error 404: No route matches this request
The first two read a document taken from a type no route produces. It is the
runtime's internal registry report — `model_provider`, `loop_engine.implemented`,
`storage.metadata.type`, `validation.checks` — so every field a reader could
reach for was `undefined` at runtime while type-checking. The third sent
`PATCH /v1/x/settings`; the route mounts `PUT`.
The SDK, therefore, is where this had to change, and the CLI is the consumer
that proves it: with the types corrected the compiler enumerated exactly what the
CLI had been reading from the wrong shape, so the two cannot move apart.
* `RuntimeSettingsSummary` and `RuntimeSettingsPatch` now describe the fields
`GET`/`PUT /v1/x/settings` actually return and accept: `revision`,
`saved_config`, `effective_config`, `restart_required`, `activation_status`,
`diagnostics`, `secret_states`, and an optional `adapters` catalogue.
`RuntimeSettingsValidationResult` is the `{valid, errors, warnings}` shape of
`POST /v1/x/settings/validate`, not the `{status, checks}` of `POST /test`.
* `settings.patch()` is a read-merge-write: read the current `revision` and
`saved_config`, merge the partial document over it, validate the result, and
write it with the revision that was read. There is no `PATCH` route to send a
partial document to, and the routes validate the whole document on write, so a
write that sent only the area being changed would be refused. A concurrent
writer is refused with `409` instead of being silently overwritten; there is no
automatic retry, because the SDK cannot know whether re-applying the patch is
still meaningful.
* Secrets survive the round trip. Reads mask a stored literal key as `********`
and the store treats that sentinel as "keep the stored value", so patching an
unrelated area cannot clobber a credential. This is asserted against the real
route rather than assumed.
* A document the runtime would reject throws `RuntimeSettingsValidationError`
before the write, carrying each issue. Measured: a fresh workspace reports
`model.api_key` / `required`, and a `${NAME}` reference the runtime cannot
resolve reports `model.api_key` / `missing_env` with the variable's name. That
detail is what the shared error path drops (#464) — without it,
`settings set-model --api-key-env TYPO` says only "Settings configuration is
invalid", which names neither the field nor the variable, and the variable is
the one thing the operator has to fix. Validate-then-write costs one extra
request and is a diagnostic, not a guarantee: `PUT` re-validates server-side.
* `settings set-model` writes a `${NAME}` reference, never a literal key, so no
secret reaches the config file or this process's arguments; `--api-key` remains
the runtime's own key. `settings validate` exits non-zero while the stored
document is invalid, so a script need not parse the output.
Verification: `npm run release:check` — 2207 passed / 15 skipped / 222 files,
typecheck (src, tests, console), build, package:check and smoke:release green.
The SDK integration test drives a real HTTP listener over real SQLite through a
recording `fetch` that delegates to the network rather than replacing it, so the
method asserted is the method served; the CLI tests drive the same listener
through the command functions and read every write back through a different
handler than the one that wrote it.
Not covered: no Docker daemon, no Kubernetes cluster, and no model provider
credentials exist on this host, so no provider boundary is exercised (a `${VAR}`
reference without a value is refused `422`, which is measured above, not
assumed). The CLI chain is also not driven by spawning the packaged binary as a
child process.
Reverse probes, each differing from the fix in one respect:
A. `patch` sends `PATCH` again -> see PR body
B. merge replaced by overwrite -> see PR body
C. pre-validation removed -> see PR body
`managed-agents workspace ...` is documented as covered in docs/api-matrix.md:88
("Create, open/register, list, resolve, and remove local workspace registry entries"),
implemented in full by src/cli/workspace-commands.ts, and ticked off in
docs/spec/tasks.md:192. The module was imported by nothing, so all five commands
answered:
error: unknown command 'workspace'
This is the fourth and last instance of that pattern (#459/#463 worker, #460/#465
session, #461/#467 environments).
This group is the one that is not an HTTP client. It reads and writes a local registry
file at $MANAGED_AGENTS_HOME/workspaces.json, so the five commands take no --port and no
--api-key, and the tests drive them against a temporary registry rather than the
developer's own. That variable was documented only in CONTRIBUTING.md; it is now in
docs/usage.md too, because the CLI is the surface that needs it.
The registration also makes a whole subsystem reachable for the first time. The orphaned
CLI module was the ONLY consumer of src/core/workspace/registry.ts, so before this the
workspace registry was dead code carrying a ticked task-list entry.
Verification:
- npm run release:check green, exit 0: typecheck across all three projects,
2191 passed / 15 skipped / 220 files, build, package:check, smoke:release.
- Every assertion is made against the registry file on disk, read directly rather than
through the module that wrote it, because there is no HTTP response to check.
- Reverse probe A, the workspace group not registered: 1 failed / 6 skipped.
- Reverse probe B, resolve not marking a miss: 1 failed / 11 skipped.
- Reverse probe C, create registering without scaffolding: 1 failed / 11 skipped.
- Reverse probe D, remove emptying the whole registry: 1 failed / 13 skipped.
Probe D exists because remove is the destructive path. It reverts the match filter to
`registry.workspaces = []`, which a single-entry test cannot detect - so the remove test
seeds two workspaces and asserts the sibling SURVIVES. A removal that dropped everything
would otherwise pass while destroying unrelated registrations.
Three behaviors are pinned that printed output alone would not catch:
- `create` scaffolds <root>/agents, <root>/skills and <root>/.managed-agents/config.yaml,
asserted ON DISK. `open` must create nothing at all, asserted by the ABSENCE of both
scaffold directories, so the two commands are not interchangeable.
- `remove` drops the registry entry and leaves every file in place. Asserted because the
same command becoming a recursive delete would destroy a user's agents and sessions
without a prompt.
- `resolve` is not a read: it bumps last_opened_at and therefore reorders the listing, so
a caller resolving in a loop silently rewrites the registry.
`resolve` and `remove` set `process.exitCode = 1` on a miss. That is process-global, so
the tests capture the incoming value in beforeEach and restore it in afterEach - a leaked
non-zero exit code would fail the run from outside the assertions.
One ordering hazard was avoided rather than slept through: entries sort by last_opened_at,
and registerWorkspace stamps new Date() at millisecond resolution, so two CLI
registrations in a row can share a timestamp and leave the order genuinely undefined. The
ordering tests seed through the registry module with explicit past times instead, so the
assertion is about the sort and nothing else.
Deliberately not in scope: no change to the registry format or its location, no --home
option (MANAGED_AGENTS_HOME is the documented override), no new workspace subcommand, and
no edit to docs/api-matrix.md - its row becomes true when the commands exist.
Closes#462
`managed-agents environments ...` is documented as covered in docs/api-matrix.md:87
("List, inspect, create, update, archive, and list worker keys"), implemented in full by
src/cli/runtime-management-commands.ts, and ticked off in docs/spec/tasks.md:105. The
module was imported by nothing, so all six commands answered:
error: unknown command 'environments'
This is the third instance of that pattern (#459/#463 worker, #460/#465 session).
Only the environments half of that module is registered. The module also holds the
`settings` group, and driving those three commands against the real API showed all three
fail:
settings get TypeError: Cannot read properties of undefined (reading 'metadata')
settings validate TypeError: result.checks is not iterable
settings set-model ManagedAgentsApiError: API error 404: No route matches this request
`GET /v1/x/settings` answers schema_version/saved_config/effective_config/adapters, not the
model_provider/storage shape the command reads; validate answers {valid, errors, warnings};
and `PATCH /v1/x/settings` is not a mounted verb (the route is PUT). Those need source
changes rather than a registration, so they stay filed as #466 and the settings group is
deliberately left unregistered instead of being made reachable while broken. The gap is
visible rather than hidden: the api-matrix and tasks rows for settings stay false.
Verification:
- npm run release:check green, exit 0: typecheck across all three projects,
2176 passed / 15 skipped / 219 files, build, package:check, smoke:release.
- Every command is driven against the REAL routes over a real HTTP listener and a real
SQLite file, and every write is read back through a different handler than the one that
wrote it, so a client-side echo cannot satisfy an assertion.
- Reverse probe A, the environments group not registered: 1 failed / 5 skipped.
- Reverse probe B, the unparseable --config-json no longer naming the option: 1 failed /
9 skipped.
- Reverse probe C, update sending a description it was not given: 1 failed / 9 skipped.
Probe C is worth recording because the FIRST version of it passed, which is a probe
failure rather than a green light. Mutating `description: opts.description` to
`?? ''` changed nothing observable, because the update route coerces an empty
description back to the stored value. A mutation the server normalises away tests
nothing, so the probe was rewritten to send `?? 'OVERWRITTEN'` and the test failed as
it should. The rule: a probe must change the bytes on the wire, not the expression in
the source.
Two behaviors had to be measured rather than assumed:
- `archive` answers with the archived row itself, so status and archived_at are on the
response the command already holds.
- Archiving is terminal: the row leaves the listing, a later single-resource read answers
404, and a second archive is refused. My first draft tried to use the second archive as
a read-back and got a 404, which is the shared `archiveResource` behavior; the test now
asserts all three consequences.
One smaller fix in the same module: parseConfigJson let the raw JSON parser error escape,
so `--config-json 'not json'` was reported as `Unexpected token 'o', "not json" is not
valid JSON` with no indication of which flag was wrong, while the sibling shape check
already said `--config-json must be a JSON object`. The two refusals of one option now
read alike, and the parser's own detail is kept.
typecheck:tests caught what vitest passed over: `app.request(...).then(...)` is typed
`Response | Promise<Response>`, so `.then` only type-checks on half the union. Fixed with
one `createEnvironment` helper instead of four call sites, which also removed the
duplication.
Deliberately not in scope: the settings group (#466); no route, SDK, or changes to
`archiveResource`; a repeat archive is asserted as measured, not changed; docs/api-matrix.md
is not edited, since its environments row becomes true when the commands exist.
Closes#461
`managed-agents session ...` is documented as covered in docs/api-matrix.md:84
("Create, message, tail, inspect, and logs"), implemented in full by
src/cli/session-commands.ts, and ticked off in docs/spec/tasks.md:102. The module was
imported by nothing, so all five commands answered:
error: unknown command 'session'
Registering them was not sufficient, and that is the second half. Driving each command
against the real routes over a real HTTP listener and a real SQLite file showed that
`session create --agent <name>` is refused:
400 {"error":{"type":"invalid_request_error","code":"invalid_agent_ref",
"message":"agent must be a standard agent id"}}
because a session's `agent` field identifies an agent by id
(src/api/routes/sessions.ts:84 requires the `agent_` prefix), and every published example
passes `$AGENT_ID` or `agent_01J8XkN5uT3vHpLqRfWdY2`. The option's help text -- which I
wrote -- claimed "name or id", a claim the server refuses. It now says id. A name cannot
simply be resolved either: migration M013 rebuilds `agents` without the `UNIQUE` that
`name` originally carried, so a name does not identify one row.
The rest was measurement rather than reading:
- `--no-stream` is pinned by value. `sessionMessageCommand` branches on
`opts.stream === false`, and that mapping cannot be read off `.opts()` before parsing:
commander applies the `--no-x` default during the parse, so a pre-parse read omits
`stream` entirely and makes a correct mapping look unverified. Measured with a
throwaway probe, which was deleted.
- `POST /v1/sessions/{id}/events` takes a batch (`events` must be an array), not one
event -- found by the seeding helper being refused.
- `POST /v1/sessions/{id}/messages` with `stream: false` persists the event and returns
`{accepted:true}`, so the message test asserts the log as well as the output: an
acknowledgment with nothing persisted would satisfy the text alone.
`session tail` is verified against the real SSE route with the log replayed from the
start. The test ends it by closing the connection, because the command is documented as
never exiting on its own, and how that forced close surfaces (a transport
`TypeError: terminated`) is deliberately NOT asserted -- a real client ends a tail with
SIGINT and pinning the transport would pin the wrong thing.
Verification:
- npm run release:check green, exit 0: typecheck across all three projects,
2165 passed / 15 skipped / 218 files, build, package:check, smoke:release.
- Reverse probe A, the session group not registered: 1 failed / 4 skipped.
- Reverse probe B, `session logs` printing the whole envelope on one line: 1 failed /
6 skipped.
- Reverse probe C, `session inspect` omitting the agent name: 1 failed / 6 skipped.
- Each probe differs from the fixed tree in exactly one respect; all three restored,
12/12 green.
The flag -> argument mapping is pinned in tests/unit/cli-program.test.ts with the command
functions replaced, and the argument -> wire behavior is pinned against the real routes in
tests/integration/session-cli-commands.test.ts. Neither alone would have caught
`worker poll` (#459), which was registered correctly and still sent the wrong body.
Boundary not covered: the chain is not also driven by spawning the packaged CLI as a
child process.
Found and filed separately: the SDK keeps only `error.message` and discards the published
envelope's machine-readable `code` and `type` (src/sdk/client.ts:327-332), so a caller
cannot branch on `invalid_agent_ref` without matching English prose (#464). The new test
asserts the code is absent, so fixing the SDK fails the test rather than passing silently.
Deliberately not in scope: no route, SDK, or session-manager change (the SDK error surface
is #464); `--agent` does not resolve names, because it cannot unambiguously; no `session
list` command, because the coverage table documents exactly five; docs/api-matrix.md is
not edited, since those rows become true when the commands exist; and the other three
orphaned CLI modules stay filed as #461 (`settings` + `environments`) and #462
(`workspace`) rather than folded in, because each has to be verified against the real
routes one at a time.
Closes#460
Make the sandbox layer able to serve three execution backends, and fix the
fail-open behavior that made backend selection unsafe.
Security fixes:
- Sandbox resolution no longer falls back to the local (unsandboxed) provider
when the requested backend is not registered. A session configured for an
isolated backend used to execute on the runtime host with no signal; it now
fails with the registered backends listed.
- Delegated sub-agents inherit the parent session's backend. They previously
hardcoded the local provider, so a Docker- or Kubernetes-configured session
still ran delegated shell commands directly on the runtime host.
- Kubernetes session Pods are created with automountServiceAccountToken false
unless an Environment names a ServiceAccount, so agent commands cannot reach
the Kubernetes API with the namespace default token.
Capabilities:
Providers declare isolatedExecution, hostFilesystem, resourceLimits, and
streamingExec instead of letting callers infer support from optional fields.
Snapshots are gated on hostFilesystem, and a capability the selected backend
lacks is reported rather than dropped: a snapshot-enabled Docker session used to
silently never snapshot, and resource limits on the local provider were ignored.
Kubernetes provider:
One Pod per session over the kubectl CLI, matching how the Docker provider uses
the docker CLI and adding no npm dependency. All cluster calls are async;
provision polls for readiness so a terminal image or config error surfaces with
the kubelet's reason instead of holding the session for the full readiness
timeout. Pod deletion failures are reported instead of leaking Pods silently.
Naming:
The registry/Settings V2 alias for the self-hosted worker (self_hosted vs
remote) had three inconsistent implementations, one of which mapped every
unrecognized value to local. It now lives in one module with two inverse
functions. Deletes the dead copy of normalizeRuntimeEnvironment.
Type checking:
The test suite was never type-checked, so a breaking interface change compiled
clean while every test call site was already wrong. Adds tsconfig.tests.json and
wires it into npm run typecheck and CI.
Removes the e2b and daytona provider types, which were selectable but had no
implementation.