798 Commits
Author SHA1 Message Date
Josh Avant 59b5fd7759 fix(qa): await completed memory replies before assertions and reset (#160929)
* fix(qa): await completed memory replies before assertions and reset

* fix(qa): restore memory config after completion timeout
2026-09-29 01:52:38 -05:00
Josh Avant 1a0e022cb1 fix(qa): correlate logical tool telemetry across Code Mode (#160973) 2026-09-29 01:27:23 -05:00
Josh Avant 0bd462d576 fix(qa): count completed image deliveries independently of progress (#160927) 2026-09-29 00:38:05 -05:00
Josh Avant 99257c0cf3 fix(qa): route directory parity checks through ls (#160920)
* fix(qa): route directory parity checks through ls

* test(qa): normalize runtime routing fixture inputs
2026-09-29 00:36:53 -05:00
Josh Avant 9433abc45a fix(qa): wait for canonical scheduling history (#160919) 2026-09-29 00:35:54 -05:00
Dallin Romney d3008b2d5e fix(qa): serialize source-backed startup proofs (#157695)
* fix(qa): serialize source-backed startup proofs

* test(qa): prove startup scripts stay serial

* test(qa): pin voice proof to mock response model
2026-09-28 19:43:05 -07:00
Peter Steinberger 78c83b538c test(release): repair release-check lanes after settle, Responses body, and scheduler changes
Release-check lanes that were also red on main asserted pre-change contracts:
- subagent settle wording now says "in this batch" (#158642)
- Responses request bodies are pre-encoded bytes (#159574); decode them
  without consuming init instead of requiring string bodies
- direct experience-review calls must bind the session MCP scheduler that
  Gateway startup binds in production (#159481)
- Docker lanes pin the plain PATH `codex` CLI for the Codex harness probe
- Telegram group-policy hot reload must name every array it replaces

Landed directly from #160746 (the native publisher needs a signed merge to
bring that branch current). Focused proof: four QA suites 200/200 in 81 s on
the merged tree; the two E2E files and credentialed live paths were not run.
2026-09-28 16:18:09 -07:00
Josh Avant a19ae3eda4 fix(qa): unblock isolated harness tool and evidence checks (#160527)
* fix(qa): recognize settled subagent batches

Share requester-settle wake recognition between mock input classification and completion handling. Accept the current batch wording while retaining installed-candidate session wording and the existing provenance check.

The real catalog-only handoff spawned and completed a child on the baseline, but QA ignored its completion and timed out. Two focused regressions failed before this fix. The real handoff now passes, all 70 owner/sibling tests pass, and check-changed passes. Standalone changed-test wall times: input 3.21s; handoff 33.76s.

* fix(qa): admit canonical repository checkpoint commands

* fix(ui): stabilize Run Inspector evidence collection

Bind rendered evidence to the public run, execution, and selected decision
receipt identities. Add a shared collector that uses the component-owned
route model and a separate page, preserving the caller's Chat surface and
unsent draft through collection and later reload.

Cover exact run/execution selection, receipt cursor reload, back navigation,
wrong requested identity, missing receipt, and page cleanup in the existing
mock-Gateway Chromium harness. Document the collector and per-tab auth setup.

The regression failed on baseline before product edits because the rendered
run identity attribute was absent. Terminal R's optional assistant transcript
idempotency key is a distinct producer gap; this does not manufacture that
key, relabel historical cells, or change product appearance.

* fix(telegram): confine QA runtime and preserve readiness evidence

Run standard-library drivers through existing Python without UV inline-script virtual environments. Confine private runtime state, drain readiness pipes, retain structural diagnostics, and require verified teardown receipts before releasing recovery state. Preserve doctor, recovery, group, and published-upgrade callers.

Validation: native sandbox regression failed before the fix; 140 harness tests pass, and 72 final focused owner/sibling tests pass in 5.73s. Build exits 0. Broad changed checks stop on two unchanged TS2459 package-update test errors. Focused lint matches all 135 baseline findings with zero additions. No live credentials or Telegram sends.

* fix(telegram): avoid native scenario barrier watchers

* fix(qa): share production publication guard contract

* fix(qa): require named message availability in mock provider

A catalog dispatcher does not advertise every delivery tool. Require the
exact message declaration and a usable invocation surface, including trusted
system/developer instruction carriers. Finish with the recovered child result
when message is absent.

Prove the absent-message regression through the mock HTTP provider and retain
catalog delivery, namespace, text-only fanout, and subagent handoff coverage.

* fix(qa): pass checkout roots to script scenarios

Expand the selected repoRoot separately from each scenario outputDir in the
maintained test-file runner. Prove the argument and CWD contract through a
real subprocess with external artifacts and fresh passing producer evidence.

* fix(qa): distinguish tool declarations from instruction prose

* fix(qa): bind forked-context evidence to native receipts

* fix(qa): bind repeated-ingress MCP scheduler

* fix(qa): add owned provider continuation checkpoints

Let maintained mock scenarios hold one session continuation before response
bytes and observe its request cursor and tool-call identity. Reuse the
provider request log and scenario signal/shutdown lifecycle; replacement
requests proceed normally without touching Gateway decision state.

* fix(qa): keep harness proofs within their owners

* test(qa): exclude all registered runtime consumers

* fix(qa): preserve runtime inputs and enforce checkpoint launches

* test(qa): skip runtime proofs when tools are absent

* fix(qa): complete checkpoint launcher test admission

* test(qa): compile repeated ingress child before execution
2026-09-28 16:48:47 -05:00
Peter Steinberger 3a300c650a fix(qa): point worktree lifecycle scenario at surviving run-end cleanup tests
#160308 deleted src/agents/worktrees/service.run-end-cleanup.test.ts, which the
managed-worktrees-workboard-lifecycle scenario still listed as a codeRef, so
extensions/qa-lab/src/scenario-catalog.test.ts failed on main. The surviving
run-end cleanup outcome coverage lives in service.test.ts (late claims, stale
lifecycle writes) and service.removal-safety.test.ts (dirty retention).
2026-09-28 06:43:20 -07:00
Peter Steinberger 39373e69f8 chore(deps): refresh dependencies through September 19 cutoff (#159401)
Refresh application, plugin, native, build and container dependencies through the fixed 2026-09-19T16:27:11Z cutoff. Migrate native TypeScript snapshot/printer APIs while preserving compilation and filesystem contracts; retain existing patches and compatibility holds. Document offline container-image preparation.

Include the verified compiler process-census and loading-clock fixture repairs and deterministic warm-history regression. Adopt the canonical production history fixes from #159924 and #159955.

Land under the maintainer's explicit approval to treat proven pre-existing CI failures as non-blocking and repair main afterward. CI36365552098 failed an unchanged Android Compose fixture's asynchronous catalog projection assertion (3518 passed,1 failed); the Android/Gradle tree matches its main parent byte-for-byte. Security and dependency reviews passed. The final rebase preserves reviewed source changes and regenerates only the intentional Node-image documentation fingerprint. See PR159401 for complete validation and the follow-up repair obligation.
2026-09-27 19:33:08 -07:00
Patrick Erichsen 71f918d0fa chore(qa): cover private operator key handoff (#157966)
* chore(qa): cover private operator key handoff

* test(qa): guard credential handoff prompt and OpenClaw config edit

* test(qa): harden operator key handoff proof

* test(qa): account for config-write migration markers in key handoff eval

* test(qa): require literal absolute secret-provider path

* test(qa): reject unusable array-root key files
2026-09-27 18:20:55 -07:00
Peter Steinberger da04ae795c fix(slack): keep top-level turns quiet and finish progress without "Working" (#159657)
* fix(slack): keep top-level turns quiet and finish progress without "Working"

Surface: Slack channel plugin progress. Requested by Peter Steinberger.

Default progress turns without a reply thread now leave only the final
answer, using the temporary hourglass reaction when typingReaction is
unset and honoring configured reactions or an explicit disable. Preserve
explicit progress presentations and native-card selection for threads.

Finish untitled Block Kit cards as Done or Failed and native summary rows
as Completed or Failed, preserving authored headlines. Remove duplicate
Working fallbacks and build both session links from the dispatched session
key with the configured session.mainKey.

Derive terminal notification and screen reader fallback text from the same
rendered blocks, including queued-turn finalization, instead of reusing
working-state text.

Document the defaults and protect delivery, reaction cleanup, terminal
titles, and session links at the Slack wire and renderer boundaries.

Keep wire-trace helpers in test support to meet the existing line cap.

* test(qa): cover quiet top-level Slack progress

* test(infra): drain migration fixture databases before cleanup

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-27 17:00:05 -07:00
Peter Steinberger a2c698f3d3 fix(slack): let verified linked admins assign sessions from Slack (#159694)
* fix(slack): resolve verified linked requester profiles

Release note: Slack users linked to an administrative profile can ask an
agent to assign a session to them from ordinary messages and app mentions
received through Socket Mode or signature-verified HTTP. Generic trusted
requester metadata now carries the canonical linked profile ID and display
label; Discord verified senders benefit from the same core path. The
sessions tool remains owner-only.

Mint transport assurance at native reception and preserve it through the
existing durable ingress kind. Keep relay and mixed-assurance batches
asserted. Preserve the exact native Slack sender ID so the host identity
handoff matches finalized SenderId instead of rejecting its lowercase form.

Resolve requester facts through the existing worker-backed identity owner.
Bind their private carrier to the admitted sender, account, and channel,
recheck context and live link authority at prompt use, and retain no new
stored identity fields. Consolidate shared message-source and slash input
types to keep existing large modules within their line-growth limits.

Validation: 199 tests passed across five serial single-file Vitest runs,
including real Socket/HTTP receiver preparation, relay replay, linked
admin/member/unlinked senders, unlink/relink, genuine-carrier transplant
rejection, and persisted sessions.assignOwner through the agent tool.
Regressions failed before their corresponding repairs. Independent Codex
review is scoped-clean through P2; git diff --check passed.

The once-authorized check-changed run stopped at line-growth violations.
Those were repaired and affected tests rerun; the one-run host limit left
that gate unrepeated and typecheck/lint unrun. No live deployment was tested.

Follow-ups: native slash commands still need host-bound requester context;
assignment for non-owner channel requesters needs a separate permission design.

* fix(channels): reuse requester facts and keep prompts stable

Forward the single prepared linked-identity result, including known absence,
through command-owner authorization instead of resolving it a second time.
Keep the existing live authority checks and owner-only sessions policy.

Release note: linked requester profiles and short assign-to-me guidance now
live in host-generated per-turn conversation info, keeping the system prompt
byte-stable across different or unlinked senders. Use the canonical account
default and preserve lowercase Slack allowFrom and stored pairing approvals
while retaining native sender ID case for identity handoff.

Correct the cross-boundary Slack test fixture to use the maintained public
artifact loader, retain the context builder overloads, return synchronous
avatar data, and await preparation after SDK socket events. No SDK export,
protocol, configuration, schema, or permission expansion is introduced.

Validation: single-file authority tests 40 passed; Slack authorization tests
57 passed, including real persisted lowercase pairing approvals and a wrong-
sender negative control. New read-count and prompt-cache regressions failed
before the repair. Affected fixture cases passed again after type/lint fixes.
Full check-changed passed core/extension and test typechecks, lint, all guards,
and its six Doctor-contract tests. Independent Codex review of the final
staged candidate is scoped-clean through P2; git diff --check passed.

* test(qa): cover verified Slack requester assignment

Prove signed Slack HTTP ingress carries a linked administrative requester
profile in per-turn user context and offers the owner-only sessions tool.
An unlinked sender receives neither fact. Seed only the personal-profile
prerequisite while the ephemeral Gateway is stopped; use public role/link RPCs.

Capture strict Crabline readiness before runtime API traffic, retain the
unaltered runtime recorder, and probe the same adapter at final capture.
The recorder regression fails before the repair and passes afterward.

The optional model-issued assignment diagnostic still fails at the separate
admitted operator-authority boundary, despite a real shared target existing.
Record this coordinator gap without broadening permissions or claiming full
assignment success. Native slash requester context remains a follow-up.

Validation: signed-ingress QA passed on Blacksmith Testbox; focused tests,
applicable lint/typecheck/guards, and independent review through P2 passed.
The one full changed-file run exposed stale installed fs-safe; restoring the
pinned version fixed core types. Remaining core-test failure is fixed on main
by e9af8ebcde. No changelog, schema, protocol, or SDK surface changes.

* test(qa): register requester profile fixture entrypoint

Declare the process-launched QA fixture as a Knip executable root so its dependency graph remains audited. The exact failing unused-file scan now passes for both production and the full tree; independent review is clean through P2.

* fix(gateway): act as the linked admin for channel turns with verified owner authority

Carry the prepared linked administrator through the existing admitted operator authority. Preserve live command-owner, role, model, and access-grant restrictions for every dispatch and active model execution. Retain the original revocation failure for shared consumers. Prove visible-session assignment and fail-closed lifecycle behavior without changing tool gating, schemas, protocol, or SDK exports.

* test(qa): prove linked Slack admins can assign agent-owned sessions

* test(gateway): preserve original operator revocation assertions

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-27 16:52:33 -07:00
Peter Steinberger 9a202d4c14 refactor(qa-lab): deslop QA Lab fourth pass (#159582)
* refactor(qa-lab): deslop QA Lab fourth pass

* test(qa-lab): defer planning selector repair to existing PR

* test(qa-lab): isolate mock request linkage coverage

* chore(qa-lab): shrink removed assertion baseline entries

* refactor(qa-lab): make scenario and dependency ownership explicit
2026-09-27 23:11:57 +00:00
Peter Steinberger ce1ca89990 refactor(tool-search): retire tool_search_code in favor of structured search and Code Mode (#159398)
* fix(tool-search): run tool_search_code in the QuickJS sandbox

Tool Search code mode spawned a Node --permission child with a node:vm
guest. Under Bun it needed an installed Node, and without one an explicit
code config silently downgraded to structured tools mode.

Run the guest through the Code Mode executor contract with the bundled
quickjs executor on every runtime. openclaw.tools.search/describe/call use
the shared namespace bridge with lazy thenables; the host admits only those
three operations through ToolSearchRuntime. codeTimeoutMs still bounds the
whole invocation, now including executor preparation. A denied or disabled
code-mode-quickjs plugin fails explicitly with next-step guidance instead of
falling back.

Remove the child source, IPC types, stderr-tail handling, and the Node and
Electron capability probe.

* refactor(tool-search): retire the tool_search_code bridge

Keep structured Tool Search and generic Code Mode as the two large-catalog surfaces. Remove the superseded JavaScript bridge and its runtime, display, and QA paths.

Doctor migrates legacy code mode to tools and removes codeTimeoutMs while preserving activation. toolSearch: true now selects structured search; JavaScript orchestration uses Code Mode exec/wait.

* test(tool-search): cover runtime behavior through structured controls

Exercise retained catalog, policy, hook, cancellation, terminal, MCP and client behavior through structured controls and their runtime owner. Delete bridge-only sandbox and JavaScript envelope cases while preserving nested call-id compatibility.

* test(tool-search): drop retired code mode prompt case

* test: repair fixtures exposed by Tool Search retirement checks

Remove the remaining retired Tool Search mode row. Preserve the session reader owner through the media retention mock and remove an unreachable queued-only branch from the ACP controls/submission fixture.

* test(upgrade-survivor): seed retired Tool Search code mode config

Author the legacy mode and timeout through every supported representative baseline CLI recipe, then require structured search and timeout removal after candidate update and Doctor. Existing config validation proves the resulting effective config.

Include the diagnostics native assignment summary in frozen target staging; the required assertion suite exposed its missing import. Node recipe and assertion tests: 238 passed, 157.87s wall. Docker validation remains with the coordinator.

* refactor(tool-search): drop the retired code-mode recovery surface

* chore: shrink assertion baseline after Tool Search retirement

* test(tool-search): drop the unused Tool Search test API

* chore: drop the retired Tool Search test API assertion baseline

* fix(e2e): drop duplicate native assignment staging line

Main now stages native-assignment-summary.mjs for frozen upgrades itself; the branch copy from the Tool Search upgrade proof became a duplicate after merging.

* build(pr): list Tool Search migration in wrapper inventory

The scripts/pr wrapper loads the Doctor config migrations at runtime, so the new Tool Search retirement migration belongs in its extracted component inventory.

* test(e2e): ship the Tool Search recipe to prepared tooling workers

Prepared tooling workers copy only listed source-relative assets, while the
upgrade survivor config recipe reads every section file by name at import.
The new tools-tool-search.json was missing, so the Docker scheduler parent
signal test's runner died with ENOENT and its polling wait reported a generic
5 s timeout that looked like a flake.

List the asset, guard the recipe directory against the preserved list, and
make the scheduler readiness wait fail with the runner's stderr once it exits.
2026-09-27 16:08:37 -07:00
Peter Steinberger 6652f7eac8 refactor: remove Tasks and TaskFlow runtime (#159179)
Remove Tasks and TaskFlow runtime, APIs, CLI, SDK surfaces and panels after the Cron, session, native execution and media completion ownership cutovers. Preserve stored rows and import provable legacy native assignments through Doctor; ambiguous ownership stays untouched with a warning.

Follows #158221, #158217, #158225, #158222, #158702 and #158776. Related: #156532. Task-specific public APIs retire immediately; retained responsibilities use their existing owners.

Maintainer-authorized administrative landing after full CI run 36312986498 attempt 2 passed on 274595e2, with subsequent actual conflicts reviewed and focused checks passing. Current PR CI preflight hits the 64 KiB changed-path metadata limit before tests (run 36335042695); its duplicate security-review status mirrors that planning failure. Review and scoped proof are recorded in the PR. Published 9.4 native import is proven; remaining native completion and 9.4 rollback witnesses are explicitly unproven.
2026-09-27 10:40:29 -07:00
Peter Steinberger cdd333e905 fix: deliver side answers while channel tasks continue (#159416)
* fix: deliver side answers while channel tasks continue

* test: refresh side-answer delivery prompt snapshots
2026-09-26 22:14:00 -07:00
Josh Avant e5bf1c0610 fix(qa): recognize instruction profile tool receipts (#158838) 2026-09-26 10:41:36 -05:00
Peter Steinberger 18b0130a21 fix(qa): follow canonical channel protocol references
Point the thread memory and personal reply scenarios at the existing SDK protocol owner after the private forwarding barrel was removed. Reproduced the missing-reference failure on clean main; all 56 catalog tests pass after the two metadata fixes. Formatting and independent P2 review pass.
2026-09-26 04:52:59 -07:00
Ayaan Zaidi 8fd72da7f2 feat(agents): let owners hand keys, config, and skill edits to their agent in chat (#158120)
## What Problem This Solves

Fixes: owners who ask their agent in chat to change a key, a config value, or one of their own skills get refused, sent to a dashboard, or told to file a Workshop proposal. Examples: "you can't post API keys here", switching the embeddings provider routed to the web-search wizard, only proposals for handwritten skills.

## User Impact

An owner can hand the agent an API key or token in chat, ask it to change config such as the embeddings provider, and have it edit skills they own. Session permission modes are unchanged: Full Access applies, restricted sessions ask.

- **Keys from chat.** Masked setup flows still keep keys out of model context and remain the default. If the user already pasted a key or token, the agent stores it in the shared secret store and points the config key at it with a `store` SecretRef instead of refusing. It never echoes the value back. The pasted message already reached the model provider and transcript; redaction covers later logs and output only, and the docs say so.
- **Existing store entries are never touched.** Each save inserts a new entry named after the config key plus a random suffix (`GATEWAY_REMOTE_TOKEN_9B139B5E231299BC`). Nothing is overwritten, revived, or deleted, and a new name can never match anything already pointing into the store, including a stale reference to a removed and purged entry. Replacing a key leaves its previous entry for `openclaw secrets store rm`. Rotating a key keeps its configured store provider alias, and the audit records the alias actually used.
- **Embeddings.** The agent now treats the memory embeddings provider, model, and key as `memory.search.*` config, not the web-search setup wizard.
- **Skills.** When the user asks, the agent edits skills they own directly: repository skill source, workspace `skills/`, project `.agents/skills/`, and configured extra skill directories. Bundled, ClawHub-installed, and plugin-provided skills are replaced by their owners' updates. For those, the agent says so and offers to capture the change as a Workshop skill.
- **Tone.** The "never request / paste credentials in chat" lines are gone from the `openclaw` tools and system-agent prompts. "Never echo secret values" stays, and so do factual pointers for flows that genuinely need a UI: channel sign-in, provider OAuth/accounts, and model onboarding.

No config option, schema, or protocol change. The Full Access permission-policy floor from the first revision moved to #158142 for its own security review.

## Why This Change Was Made

After #149870, approved config writes may target any path. What still blocked owners was model-facing text telling the agent to refuse credentials, plus the missing ability to store a chat-provided value anywhere but plaintext config.

`config_set_ref` gains an optional `secret` argument (read without trimming; only emptiness is checked). With it, the system agent:

1. registers the value for redaction when the proposal is built;
2. keeps the key's existing store provider alias when it has one;
3. sends one `secrets.writeForConfigRef` command to the SQLite state worker with the requester's live-authority guard. The host re-checks that guard at the worker's transaction and commit admission (`createSqliteWorkerWriteAdmission`), so a run stopped while the command is queued writes nothing. The transaction inserts a new row under a freshly minted `NAME_<16 random hex>`;
4. writes the ref through the existing config writer, which re-checks authority. If that write fails (before or after the writer commits), OpenClaw rereads the config and the error says the key was saved as `<NAME>` and whether the config key points at it. There is no automatic delete: another consumer may have linked the fresh entry, or the writer may have committed before failing;
5. the normal config reload picks up the new ref, since its id always changes.

Nothing new runs SQLite on the Gateway main thread.

<details>
<summary>Out of scope / follow-ups</summary>

- Found while proving this: in Full Access, after a delegated change applies, the next agent turn in the same chat fails with `SQLite database already belongs to another worker backend`. It reproduces on unmodified `origin/main` (`71bb516`) with a `logging.level` change followed by one more message. This PR does not fix it.
- Built-in provider sign-in and model onboarding stay handoffs; they own live verification of the active inference route.
- Other secret-store set/delete paths remain synchronous migration debt, as `worker-access.md` already records.

</details>

## Evidence

Real Telegram Test Server (Convex-leased userbot, fresh Gateway, QA mock provider, Full Access, tester is owner), first revision:

| | Screenshot |
|---|---|
| Token given in chat, applied with no approval prompt and no refusal (synthetic QA token) | ![User sends a remote Gateway token and asks to save it; the agent replies without refusing or asking for approval](https://github.com/user-attachments/assets/4f68f3db-225c-4047-984d-ed0cff79cd0c) |

### Final effects at this head

qa-channel scenario `system-agent-owner-trust` passes through a real Gateway and state worker. The Gateway is seeded with an unrelated `GATEWAY_REMOTE_TOKEN` entry, then:

1. A command-allowed **non-owner** (`bob`) sends the key. The `openclaw` tool is owner-only.
2. The **owner** (`alice`, Full Access) sends it.

Captured step details (redacted by the Gateway; the store ref id prints as `__OPENCLAW_REDACTED__`):

```json
{
  "nonOwnerEntryNames": ["GATEWAY_REMOTE_TOKEN"],
  "storedRef": { "source": "store", "provider": "default", "id": "__OPENCLAW_REDACTED__" },
  "storeEntryNames": ["GATEWAY_REMOTE_TOKEN", "GATEWAY_REMOTE_TOKEN_9B139B5E231299BC"]
}
```

- After the non-owner turn: only the seeded entry exists and `gateway.remote.token` is unset.
- After the owner turn: `gateway.remote.token` is a `store` SecretRef, the token is in its own minted entry (`GATEWAY_REMOTE_TOKEN_9B139B5E231299BC`), and the seeded entry's `updatedAt`/`updatedBy` are unchanged. No approval prompt was posted, and the token is absent from chat, config, and the store listing.

**Revoked request**, through the production worker (Node main thread, real broker, `writeSecretStoreEntryForConfigRef`). The requester's guard passes the caller's check, then reports the run stopped:

```text
seeded: [ 'GATEWAY_REMOTE_TOKEN (cli)' ]
revoked request rejected: requesting run is no longer active
after revoked request: [ 'GATEWAY_REMOTE_TOKEN (cli)' ]
owner request saved as: GATEWAY_REMOTE_TOKEN_F1657B971F691824
after owner request: [ 'GATEWAY_REMOTE_TOKEN (cli)', 'GATEWAY_REMOTE_TOKEN_F1657B971F691824 (openclaw)' ]
seeded value intact: true
```

Tests (each fails without the behavior it covers):
- production worker path (`secret-store-config-ref.worker.test.ts`, forked database-worker lane with the real broker): a chat secret gets its own minted entry beside a live `GATEWAY_REMOTE_TOKEN` without touching it; a requester revoked after the caller's check writes nothing;
- store kernel: a refusal at commit admission rolls the transaction back; each save mints a new `NAME_<hex>` and leaves the key's previous entry unchanged; a stale name whose entry was removed and purged still resolves to nothing after a chat save;
- operations: stored and referenced with no value in output or audit; authority gone before the store write writes nothing; a failed config write names the saved entry and leaves it in place; rotating a key keeps its configured store provider alias, and the audit records it;
- tool: proposes a store write without repeating the key, preserving leading and trailing whitespace.

Measured single-worker wall time per new or materially changed test file at this head (`node scripts/run-vitest.mjs run <file>`, local M-series; vitest Duration includes import and setup):

| File | Tests | Wall | Vitest duration |
|---|---:|---:|---:|
| `src/secrets/store/secret-store-config-ref.worker.test.ts` (new, database-worker lane) | 2 | 14 s | 2.05 s |
| `src/secrets/store/secret-store.test.ts` | 32 | 16 s | 13.37 s |
| `src/system-agent/operations.test.ts` | 43 | 18 s | 15.30 s |
| `src/agents/tools/system-agent-tool.test.ts` | 36 | 15 s | 12.89 s |

QA scenarios: `system-agent-owner-trust` (mock-openai) runs in about 27 s after build; `skill-owner-direct-edit-live` is live-frontier only and took about 3 min with `claude-cli/claude-sonnet-4-6`.

Wording pins for the removed lecture text were deleted. The focused store, worker, exclusivity, operations, tool, approval, and delegate suites pass. `node scripts/check-changed.mjs` passes every gate except core lint, which fails only on three files this PR does not touch (`server-chat-metadata-lifecycle.integration.test.ts`, `session-companion-ask.ts`, `app-sidebar-session-list-render.ts` over `max-lines` on the base); oxlint on the changed files is clean.

Security decision: a Full Access owner's pasted key goes to the Gateway-wide team store without a separate approval. Maintainer (@obviyus) accepted this in the PR conversation.

**Rotation with a second consumer**, through the production worker (Node main thread, real broker). A second consumer references the key's first entry before the next save lands; value fingerprints only:

```text
owner saves key #1 -> MODELS_PROVIDERS_OPENAI_API_KEY_6D96F6E92B59E935 (sha256:4a5c5a4aa8de)
second consumer now references MODELS_PROVIDERS_OPENAI_API_KEY_6D96F6E92B59E935 (e.g. linked while the next save is queued)
owner saves key #2 -> MODELS_PROVIDERS_OPENAI_API_KEY_4066F18ABAAE6903 (sha256:28bc4e3fe10d)
second consumer's entry MODELS_PROVIDERS_OPENAI_API_KEY_6D96F6E92B59E935 after rotation: sha256:4a5c5a4aa8de
unchanged: true
```

**Config write fails after the save, with a second consumer on the fresh entry**, through the production worker and the system-agent apply path (fingerprints only):

```text
owner result: Saved the secret as GATEWAY_REMOTE_TOKEN_6B68C031E571F81C, but could not point gateway.remote.token at it: config write failed after commit (rollbackStatus: not-restored). Retry, or remove the entry with `openclaw secrets store rm GATEWAY_REMOTE_TOKEN_6B68C031E571F81C`.
config gateway.remote.token: null
second consumer's entry GATEWAY_REMOTE_TOKEN_6B68C031E571F81C: sha256:09ae5b4fd36b
second consumer keeps the credential: true
```

**Stale reference to a removed and purged entry**, production worker for purge and save:

```text
purged rows: 1
stale ref GATEWAY_REMOTE_TOKEN after purge: SECRET_STORE_NOT_FOUND
chat save -> GATEWAY_REMOTE_TOKEN_8F698B5396990ADD resolves (value hidden)
stale ref GATEWAY_REMOTE_TOKEN after chat save: SECRET_STORE_NOT_FOUND
```

**Owned-skill edit with a live model.** New scenario `skill-owner-direct-edit-live` (live-frontier; run with `claude-cli/claude-sonnet-4-6`, subscription auth) passes at this head. It seeds workspace skill `qa-owner-greeting` replying `OWNER-GREETING-V1`, and the owner asks in plain words: "Please change my qa-owner-greeting skill so it replies OWNER-GREETING-V2 instead of OWNER-GREETING-V1." Captured result:

Skill file after the turn:

```markdown
---
name: qa-owner-greeting
description: Greets the owner with a fixed marker
---
When the user asks for the owner greeting, reply with exactly: OWNER-GREETING-V2
```

Agent reply: "Let me find the skill file. Done. The `qa-owner-greeting` skill now replies `OWNER-GREETING-V2` instead of `OWNER-GREETING-V1`."

The model edited the skill file in place and confirmed it; it did not refuse or file a Workshop proposal.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-26 16:20:43 +05:30
Josh Avant 0e31c78efe fix(qa): recognize canonical nested tool receipts (#158641) 2026-09-26 04:02:09 -05:00
Josh Avant ca5f05c87e fix(qa): use bounded instruction injection evidence (#158634) 2026-09-26 03:55:03 -05:00
Josh Avant a8272f099f fix(qa): use owner for approval resolver fixture (#158622) 2026-09-26 03:54:48 -05:00
Josh Avant 6c9e41fdca fix(qa): restore config apply restart wake-up proof (#158608) 2026-09-26 03:54:43 -05:00
Peter Steinberger 3c449bd35b refactor(cron): read run history through its worker (#158222)
* refactor(cron): read run history through its worker

* refactor(cron): clean up retired history type exports

* test(cron): finish history type and ratchet cleanup

* test(cron): prove published upgrade history reads

* test(cron): retain canonical failover type contract

* ci: refresh Cron checks after main import repair

Re-evaluate the unchanged reader against main after 7c4866c73b repaired the Team Reports test import. Production and upgrade-harness source remain unchanged.

* fix(cron): complete wrapper cutover and upgrade diagnostics

* test(cron): complete source references and delivery-free upgrade seed

* test(codex): keep catalog progress on one fake clock

* test(cron): align invalid run id assertion with owner

* ci: refresh Cron checks after async task main repair

* ci: refresh Cron checks after main QA reference repair

* ci: refresh Cron checks after Linux fixture repair

* ci: refresh Cron checks after workspace quota repair

* ci: refresh Cron checks after main progress type repair
2026-09-26 00:07:06 -07:00
a0710b4377 fix(qa): accept native Codex truncation in cache evidence (#158454)
Recognize positive-count native history truncation markers in the correlated tool output while retaining read and follow-up evidence checks.

Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: RomneyDa <6581799+RomneyDa@users.noreply.github.com>
2026-09-25 22:37:42 -07:00
Peter Steinberger a99102804e refactor: share filesystem admission and cleanup with fs-safe (#158313)
* refactor: share filesystem admission, walks, and cleanup with fs-safe

* style: satisfy lint in fs-safe adoption changes

* test(daemon): map fs-safe reads into launchd fixtures

Keep the real fs-safe reader and its descriptor admission on the fixture backing files. Load fixture mocks before filesystem consumers and move shared state out of the oversized launchd suite.

* test(gateway): assert missing transferred manifest by errno

* fix(qa): drop retired Telegram formatter reference

The formatter refactor in #158164 folded format-render.ts into format.ts. Keep the existing format.ts code reference and remove the obsolete path so the scenario catalog reference check follows the current owner.

* fix(qa): preserve exact worktree cleanup identity

Keep the original bigint device and inode comparison when replacing the local forwarding helper. The general fs-safe matcher tolerates unknown Windows identity fields and cannot authorize destructive cleanup. Four real-directory receipt regressions fail the previous candidate; 35 cleanup/runtime cases and changed checks pass, with clean independent review.
2026-09-26 01:58:46 +00:00
Ayaan Zaidi a71e5566d7 fix(approvals): let OpenClaw change approvals complete in chat (#157947)
## What Problem This Solves

Fixes: when an agent asks OpenClaw to change config, restart, or manage plugins/channels/agents, the approval never reaches chats without a native approval card, and `/approve <id>` rejects it. The agent's tool call then blocks until the 10-minute expiry, and users are sent to the Control UI or a terminal to approve a change they asked for in chat.

## User Impact

User impact: every approval an agent needs can be completed in the chat where it was requested.
- Native approval cards (Telegram, Discord, Slack, WhatsApp, Signal, Matrix, Teams, iMessage, Google Chat) keep owning their chats.
- Every other requesting chat now gets the change summary with **Allow once** / **Deny** buttons where the channel renders them, plus a `/approve <id> allow-once|deny` line.
- `/approve` now resolves OpenClaw change approvals, not only exec and plugin approvals.

No config options, protocol, or schema changes.

## Why This Change Was Made

Delegated OpenClaw changes ("system-agent" approvals) already had native cards (#134670). Two gaps remained outside those cards:

- **Request delivery.** The request was created with delivery turned off, so the shared approval forwarder never posted a fallback. The forwarder now has a system-agent strategy that always targets the requesting chat and is suppressed whenever that chat's native card handles the request. It uses the same typed-button payload and resolved/expired messages as exec and plugin approvals. The system-agent owner publishes the outcome of a decision (applied, or denied) once; approval publication only adds expiry and cancellation, so a denial produces one chat update. The fallback answers only the live messaging chat that made the request: terminal and Webchat requests never fall back to the session's saved chat, and expiry is reported once, from the Gateway's recorded expiry rather than a local chat timer, so a change approved just before the deadline and applied after it reports its applied outcome instead of a false expiry.
- **`/approve`.** The command probed only exec and plugin approvals. It now also resolves system-agent approvals through the canonical `approval.resolve`, after confirming the id is in `openclaw.approval.list`. Canonical resolution records a kind mismatch as a deny, so `/approve` confirms the owner first instead of probing. On channels with their own approver settings (Telegram, Discord, Slack, …), the Gateway checks the sender as the reviewer, the same check as native buttons. Everywhere else only a configured owner (`commands.ownerAllowFrom`) can approve an OpenClaw change: `/approve` sends the sender as the reviewer, and the Gateway checks owner custody against the current config inside the approval store's final decision guard. An ordinary command-authorized sender is refused, and an owner removed after sending `/approve` cannot complete the decision.

Unchanged: approval authority stays bound to the requesting run; free-text "yes" never approves; Full Access still auto-applies; Control UI and the apps can still decide.

Agent-facing text now says what happens: for runs from messaging channels, the `openclaw` tool description says the change waits for approval in this chat (buttons or `/approve`). Webchat and terminal runs, which the chat fallback cannot reach, are told to approve in the Control UI or OpenClaw apps. The gateway-only prompt line points to `openclaw` and `/restart` instead of "ask human".

## Evidence

Real Telegram (Test Server userbot, Convex-leased credentials, fresh Gateway, mock provider), restricted `exec` mode, agent asks OpenClaw to `set logging.level "info"`:

| Run | Setup | What the user saw | Tap | Result |
|---|---|---|---|---|
| A | tester is owner (native cards on) | 🔒 native approval card, no `/approve` text | **Allow Once** | `answerCallbackQuery` + 2× `editMessageText`: "approved. Applying" → "approved and applied"; final reply delivered |
| C | tester is owner, `channels.telegram.execApprovals.enabled: false` | fallback message with change summary, `/approve <id> allow-once\|deny`, and **Allow Once** button | **Allow Once** | message edited to "approved. Applying…", then "approved and applied" posted; Gateway config on disk has `logging.level: "info"`; final reply delivered |

Telegram Web screenshots from the same leased test account (cropped to the conversation):

| | Pending | After **Allow Once** |
|---|---|---|
| Native card (A) | ![Native Telegram approval card with Allow Once and Deny buttons](https://github.com/user-attachments/assets/5d3de39a-339c-4b7b-80d5-84f5d46cc575) | ![Native card edited to approved and applied, followed by the final reply](https://github.com/user-attachments/assets/3e993ffe-3f85-4941-9f31-0904719ea608) |
| Fallback message (C) | ![Fallback approval message with change summary, /approve line, and Allow Once and Deny buttons](https://github.com/user-attachments/assets/ca78d61c-627e-47b0-9505-66fa7162e540) | ![Fallback message resolved: approved and applying, then approved and applied, then the final reply](https://github.com/user-attachments/assets/fb4b5e28-074b-4f69-bbcb-053398b94c25) |

With no owner configured, the owner-only `openclaw` tool isn't exposed, so no change is proposed (config unchanged). That's the existing design; DM pairing sets the first owner.

qa-channel (no native approval cards), new scenario `system-agent-chat-approval`: the agent's delegated change posts the approval in the requesting conversation. A command-authorized non-owner (`commands.allowFrom` includes them, `ownerAllowFrom` does not) sends `/approve <id> allow-once` and gets "❌ Only the owner can approve OpenClaw changes in this chat."; `logging.level` is unchanged and the approval stays pending. The owner's `/approve <id> allow-once` then resolves it; `logging.level` becomes `info`; exactly one final reply. On `origin/main` the same scenario times out waiting for the approval in the chat.

Tests: `/approve` (resolves a pending OpenClaw change canonically; never submits a canonical decision for an id owned by another kind; on a channel without approver settings, a non-owner and a revoked owner never reach the canonical decision), channel custody (without approver settings only a configured owner holds custody, and only for OpenClaw changes), `approval.resolve` (custody revoked between the request and the final write leaves the approval pending), approval publication (allowed/denied changes leave the chat outcome to the system-agent owner; expired/cancelled are published once), forwarder (requesting chat gets the prompt and outcome without `approvals.*` config; a running native card suppresses it; terminal and Webchat requests never reach the saved session chat; a change applied after the deadline reports its outcome with no false expiry; the recorded expiry is reported once), plus existing gateway approval and system-agent suites. The new owner-custody and single-publisher tests fail with the fix reverted. Shared approval and forwarder fixtures moved to sibling `*.test-support.ts` modules so the test files stay under the line cap. Wall time for the touched suites: `commands-approve` + `approval-publication` + `exec-approval-forwarder` + `system-agent-approval` + `approval` ran in 59.5s across 3 Vitest shards. `pnpm tsgo:core` clean.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-25 15:05:57 +05:30
Peter Steinberger 2abecd703d chore(deps): refresh dependencies with a seven-day cutoff (#157238)
* chore(deps): refresh dependencies with a seven-day cutoff

* fix(deps): preserve Teams and jsdom integration contracts

Use the Teams SDK public token and processing APIs while keeping SSO sender
checks ahead of native token operations. Remove obsolete ambient declarations
and route workarounds, and cover the SDK routing with real processing tests.

Adapt the test environment to jsdom private-field bindings, preserve file bytes
and registry cleanup, and preload it through native Node and Bun workers.

* fix(test): preserve jsdom window and fixture contracts

* fix(ci): keep typecheck cache reuse within matching inputs
2026-09-25 02:38:45 +00:00
Peter Steinberger 809808b76a feat: show live cloud worker pool in settings (#157200)
* feat: show ready cloud worker pool in settings

* test: complete cloud worker cleanup fixtures

* test: wait for sidebar before plugin registration check

* test: run composed gateway scenarios in source lanes

* fix(gateway): opt in to live-authorized pool details

* test: document native TypeScript shutdown noise workaround
2026-09-24 13:42:58 -07:00
Peter Steinberger 6cf4b810c0 test(qa): select compaction requests by their session identity
After mock request ownership moved to transport affinity, the mutating-tool
compaction scenario still searched prompt text for the session identifier.
The provider recorded the real context overflow, but those stale selectors
found no owning requests. A copied foreign identifier in prompt content
could also match the wrong session.

Select overflow, write, and continuation requests using their existing
request.sessionId field. Exercise the actual catalog expressions with the
owner identifier absent from prompt text and a foreign identifier copied
into it. Preserve the mutation, pruning, ordering, count, and failure checks.
No shared helper or production session contract changes.

The real scenario failed before the selector repair and passed all three
after samples: request size fell from 333954 to 119338 bytes, with exactly
one logical write, one authenticated wire success, and one compaction.

The changed catalog test passed 20 pressure runs at eight workers on two
CPUs plus a CPU contender. All 58 catalog consumers passed across the
four canonical CI shards (312 files, 4485 cases); the affected 78-file,
1062-case shard passed three times. Literal one-worker unit cost: 2.794s.
A fresh real scenario replay on main, including the guarded session
observer, passed in 21.309s. Proof runs: 35954197301 and 35957749551.

The changed checker completed all type graphs, guards and dead-code scans.
Remaining global lint was stopped in favor of scoped type-aware lint with
a rejecting negative canary; that substitute, import-cycle checks,
formatting, whitespace checks and independent P2 review passed.
2026-09-23 22:08:14 -07:00
Josh Avant a564482f1e fix(qa): preserve WhatsApp silence observation windows (#156734) 2026-09-23 23:37:15 -05:00
Peter Steinberger 6386dc4665 fix(qa): resolve mock sessions from transport affinity
Removing sessionId from the Runtime prompt intentionally stabilized the
cached prefix, but the QA mock still parsed it. Requests then shared
anonymous scenario state and terminal subagent settlement lost its
requester identity.

Observe full harness-generated ids through the existing before-run hook.
Resolve existing transport affinity by exact id, otherwise a unique
64-character prefix; reject unknown, ambiguous, or missing scoped ids.
Keep scenario state per session and expose canonical identity to the
subagent-completion scenario. Preserve affinity through QA proxies and
omit HTTP-only compatibility from the native harness catalog projection.

Update the settlement fixture for the published sessions.list contract
and align the package fixture with candidate-declared schema versions
and the existing explicit immediate-drain option.
2026-09-23 20:51:53 -07:00
Peter Steinberger 7dbfab8c2c chore(deps): refresh dependencies with a seven-day cutoff (#156363)
* chore(deps): refresh dependencies with a seven-day cutoff

* fix(deps): preserve runtime and tooling contracts after upgrades

* fix(deps): align fixture lifetimes and preserve Unicode contracts

* fix(ui): publish goal rejections and align integration fixtures

Publish rejected goal actions through the existing renderer lifecycle. Start the Session Share service in integration fixtures, preserve same-job workflow authority, and distinguish pnpm launcher escalation from detached cleanup ownership.

* test: follow Corepack ownership and await navigation handoff
2026-09-23 17:35:48 +00:00
Peter Steinberger 6dc9d86447 fix(qa): release terminal children after requester settlement
The mock released child results when the parent's HTTP response was sent,
before the parent execution owner closed. Timestamp-prefixed all-settled
inputs also missed the fixture matcher and produced a generic response.

Correlate pending children with the exact runtime parent and use the existing
scenario wait loops to verify sessions.list reports that parent done, inactive
and not aborted. Keep the existing budgets and every direct-delivery, exact-send,
privacy and restart assertion. Recognize timestamped settlement inputs without
accepting quoted history. Migrate all mock-server and scenario callers together.

The six changed standalone test files each pass 20 full runs, with 91 focused
Gateway/QA cases and 43 HTTP sibling cases passing. Linux original profile 4
passes three times (27/27 scenarios, zero skipped), plus native Telegram and
QA-channel/empty flows. Actions proof: 35846151730. Single-worker file walls:
parser 2.47s, gate 2.48s, handoff 11.67s, routing 15.33s, surface 24.26s.

Types, scoped type-aware lint with a negative canary, formatting and P2 review
pass. Full lint declaration preparation hit ancestor-install isolation and used
the requested scoped substitute. Installed-package upgrade/rollback was not run
because no candidate tarball was configured; its migrated callbacks typecheck.
2026-09-23 04:15:55 -07:00
Dallin Romney 006dfc4527 test(qa): catalog Mattermost delivery custody (#155470) 2026-09-23 02:54:30 -07:00
Peter Steinberger 3e4ab6bed8 refactor: retire pre-June import and verification compatibility (#156285)
* refactor: retire pre-June import and verification compatibility

Remove pre-June task, flow, and plugin-state sidecar imports, obsolete
runtime chunks, package/installer validation exceptions, the old MCP
attachment fallback, and the April self-upgrade lane with its orphan helpers.

Leave retired data files untouched and document migration through 2026.6.1.
Preserve June-and-later contracts and September delivery recovery receipts.

Refs #156190

* docs: route legacy upgrades through 2026.9.5

* test: await Telegram fixture lifecycle events

Replace the setup stopwatch with the actual stop event or terminal run outcome. Keep cancellation assertions and outer execution bounds, and prove early terminal outcomes fail promptly.
2026-09-23 02:22:22 -07:00
Vincent KocandDallin Romney ca78e11c09 fix(qa): stabilize runtime tool evidence (#120353)
* fix(qa): wait for native patch transcript completion

Punchcard-Session: golden-valley-workshop-br

* fix(qa): make sessions spawn fixtures deterministic

Punchcard-Session: golden-valley-workshop-br

* fix(qa): bound sessions spawn evidence

---------

Co-authored-by: Dallin Romney <dallinromney@gmail.com>
2026-09-23 01:37:27 -07:00
Dallin Romney 783c4d868f chore(qa): catalog Teams restart quote delivery (#155472)
* test(qa): catalog Teams restart quote delivery

* test(qa): align Teams scenario with file scope
2026-09-22 18:48:59 -07:00
Josh Avant fdae00fe82 chore(qa): qualify execution identity across live boundaries (#156033)
* test(qa): repair forced restart and message inspection harnesses

* test(e2e): qualify installed execution identity persistence

Extend the packed npm onboarding owner with audit opt-in, one deterministic local turn, installed CLI inspection before and after Gateway restart, and isolated-state/privacy assertions.

Qualification from 267bb9d2: new shell-flow proof fails before the harness change at the missing opt-in. All 53 focused support tests pass; final changed shell/SQLite selection took 26.94s with one worker. Changed checks and P0-P2 autoreview pass; the full export scan was reused after brace-only lint fixes. Packed Docker proof remains a remote handoff.

* test(qa): qualify live execution and channel identity

* test(audit): qualify execution identity lifecycle gaps

* test(qa): align suppression audit with required replies

* test(qa): qualify Telegram participant identity

* test(qa): provision Telegram identity fixtures

* test(qa): qualify private-production Telegram identity

* test(e2e): preserve installed execution identity proof

* test(qa): await post-delivery memory maintenance

* fix(qa): repair qualification CI boundaries
2026-09-22 20:15:02 -05:00
Peter Steinberger ddde193e26 fix(qa): key Code Mode terminal evidence on the exec surface (#155811)
## Summary

Unblocks the 2026.9.6 Full Release Validation (Release Checks run 35743792326, jobs 106800921622 and 106800921540): `compaction-retry-mutating-tool` failed identically in the parity (candidate) and runtime-pair (core) lanes with `Code Mode terminal continuation did not report successful completion`.

## Root cause

The scenario picked the expected terminal-continuation shape by provider variant: `openai` meant the Codex-native `Script completed\n...` text and `anthropic` meant the guest JSON `{ "status": "completed", ... }` result. The `Script completed` text only exists in the Codex harness (`extensions/codex`); the OpenClaw runtime's Code Mode always returns the guest JSON regardless of model.

Until #155614 an absent `tools.codeMode` defaulted to off, so this mock-openai lane performed the write through the direct `write` tool and the `writeWireToolName !== 'exec'` guard short-circuited the assertion. #155614 made the absent setting behave as `"auto"`, so `openai/gpt-5.5` now routes the write through guest `exec` and the never-satisfiable native branch fired. The product behaved correctly: the local repro shows the terminal continuation carrying `{"status":"completed", ..., "value":{"changed":true,"created":true,...}}`, the file with the exact expected content, and one compaction.

## Fix

- Mock OpenAI records the resolved Code Mode exec surface (`native` = freeform custom tool, `guest` = `code`-schema tool) on every request snapshot (`codeModeExecSurface`).
- The scenario keys the terminal-evidence assertion on that surface instead of the provider variant, still failing closed when neither surface is present. This is strictly stronger for the OpenClaw runtime: OpenAI-model guest runs are now actually verified for `status === "completed"` instead of being skipped.
- Catalog test pins updated to the new discriminator.

## Proof

- `node scripts/run-node.mjs qa suite --provider-mode mock-openai --parity-pack agentic --concurrency 1 --model openai/gpt-5.5 --alt-model openai/gpt-5.6-luna-alt --scenario compaction-retry-mutating-tool` on origin/main: fail (reproduced CI). With this change: pass, details `wireTool=exec ... wireSuccesses=1 compactions=1`.
- `scenario-catalog-compaction`, `scenario-catalog`, `mock-openai/server`, `agentic-parity-report` tests: 417 passed.

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-22 10:01:58 -07:00
35c3385daf fix(recovery): preserve restart-safe tool ownership (#137483)
* fix(qa): bind restart checkpoints to pending waits

* fix(sessions): preserve restart recovery safety guard

* fix(qa): honor current restart tool declarations

* fix(qa): honor unsafe restart tool declarations

* fix(qa): interrupt restart checkpoints before they drain

* fix(recovery): reject restart status before final capture

* test: align hot reload publication expectations

Reuse the exact three-line prerequisite from steipete PR155428 at 172e0aec90f44445ff24a4f1ae95af20993597c2. Full-file postimage and unchanged production contract match the donor; its four failing-case and six owner-supersession controls are reused. Supplemental API-direct full-file review is P0-P2 scoped-clean. No product behavior changes or assertion weakening.

* fix(models): retire catalog resources before closing work scope

Keep request cleanup admitted after credential work settles; drain the owner after discovery registry retirement. Preserve the existing cancellation barrier and normalize the equivalent Discord fixture to landed main.

* fix(qa): sanitize Slack failure causes without lint suppressions

Repair the production lint-suppression inventory failure from three
unlisted Slack QA directives. The existing failure sanitizer now creates
safe errors containing only validated Slack codes/scopes, preserving the
messages while dropping raw SDK headers and nested causes. Reuse the
existing Slack sender and retain exact channel receipt validation.

Validation: original lint inventory test reproduced the failure; repaired
inventory 4/4 and Slack owner suite 17/17 pass (28.33s combined wall with
one worker; Slack suite 9.14s total, about 0.7s test execution). Full
check:changed and independent P0-P2 review pass. No lint baseline changes.

(cherry picked from commit 44559eb315)

* test(qa): reconcile Slack runtime fixture with native send helper

Preserve the real Slack send helper in the adapter partial mock so deadline, cancellation, receipt and cleanup coverage executes native requests. Remove only the three obsolete inventory rows after landed sanitizer commit 44559eb315 removed their directives.

---------

Co-authored-by: Jason (Json) <263060202+fuller-stack-dev@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-22 06:01:07 +00:00
Dallin Romney 2583b9d937 test(qa): register repeated-request recovery (#154694)
* test(qa): register repeated-request recovery

* test(qa): remove ignored vitest timeout
2026-09-21 21:13:30 -07:00
Dallin Romney 3312f47832 test(qa): register visible child waited-send proof (#154642) 2026-09-21 21:13:15 -07:00
Dallin Romney 3ad11991d5 test(qa): register cron startup recovery e2e (#154601)
* test(qa): register cron startup recovery e2e

* docs(qa): point cron scenario at runtime startup owner
2026-09-21 21:12:43 -07:00
Dallin Romney cfe7a64a76 test(qa): register heartbeat session routing e2e (#154550)
* test(qa): register heartbeat session routing e2e

* docs(qa): narrow heartbeat restart coverage claims
2026-09-21 21:12:28 -07:00
Ayaan Zaidi 43ed8cee3f feat(qa): add Convex-backed Discord and Slack E2E skills (#153471)
Related: #153456

## What Problem This Solves

Agents need reusable Discord and Slack E2E workflows with the same Convex-login setup as the Telegram userbot skill, rather than separate secret setup and ad hoc channel probes.

## User Impact

Developer-only: repository skills provide same-lease readiness, reusable YAML flows, native message/file/thread actions, correlated Gateway replies, and explicit evidence/cleanup instructions. Telegram reuses the extracted broker-discovery owner. Curated QA defaults and production channel behavior stay unchanged.

The dedicated Slack QA Driver now has the operator-approved reaction/file read/write scopes and was reinstalled. Its existing pooled token remains valid; no workspace-admin privileges, user-token scopes, or production-app changes were needed. The setup manifest includes those scopes.

## Why This Change Was Made

The skills extend the existing QA Lab transport, Gateway, scenario, and credential-lease owners instead of adding another runner. Opt-in `--doctor` and repeatable `--scenario-file` use Convex CI leases and deterministic `mock-openai` by default, while preserving explicit overrides.

Native-write flows do not retry ambiguous writes. Cleanup retains lease authority through Gateway shutdown, captures final native receipts before temporary state removal, and removes only owned fixtures. Native and cleanup Slack requests have bounded settlement with no write replay. Unanswered Gateway mutations remain explicit uncertainty and preserve runtime evidence instead of disappearing. Discord recording remains active through shutdown; threads are archived rather than deleted when the lease lacks thread-management permission.

Native API receipts are not model-tool or visual proof. Human slash commands, component clicks, modals, ephemeral replies, and rendering retain explicit client-testing requirements. Broader testing capabilities are deferred to follow-up PRs.

## Evidence

- Real Discord full native lifecycle passed on `ea1b0536204`: messages/replies/pagination/edit, reaction add/remove, public thread, attachment upload/delete, a fresh correlated SUT reply, continuous recording, and successful owned cleanup.
- Real Slack full native lifecycle passed on `db88f0ac1cc`: all five steps passed, including quiet ingress, threaded Gateway reply, stored edit/delete and reaction add/remove verification, and exact file upload/readback. Cleanup reported zero remaining owned messages/files/reactions and zero pending, uncertain, or failed operations.
- Both native runs used real transports, a temporary Gateway, `mock-openai`, and Convex CLI discovery with both broker environment variables unset. Telegram/shared discovery regressions: 30 passed.
- Earlier focused QA set: 301 passed; subsequent affected-consumer set: 298 passed (overlapping, not additive). Extension typecheck, changed-TypeScript type-aware lint, source ratchets, docs checks, and whitespace checks passed during preparation.
- Landing reconciled one pinned main snapshot, `eac221ae78bd`, preserving both lifecycle guards. After restoring its frozen dependency versions, the merged Gateway lifecycle, Slack ownership/adapter, and scenario deadline suites passed: 4 files, 52 tests. Earlier native artifacts are retained as pre-reconciliation proof, not relabeled as executions of the merge head.
- Repaired-head Slack lifecycle passed on `62ed1326be41`: all five steps, reaction/file readback, final receipt capture, and clean owned teardown. This rerun covers the bounded-request and uncertain-capture repairs, on the reconciled base.
- Both teardown regression controls failed on the pre-repair implementations (missing HTTP deadline; unanswered mutations dropped). The repaired owner/consumer set passed all 73 tests across six files. Targeted type-aware lint, extension typecheck, docs MDX and whitespace checks also passed.
- Hosted CI will not be awaited, per the operator's explicit landing instruction; no full CI success is claimed.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-22 09:03:22 +05:30
Peter Steinberger d0f40a8de2 chore(deps): refresh dependencies with seven-day cutoff (#154652)
* chore(deps): refresh dependencies with seven-day cutoff

Advance eligible runtime, native, release, and development dependencies published by 2026-09-14T07:00:00Z. Preserve compatibility holds and existing reviewed newer pins. Synchronize release integrity checks and scoped overrides; remove the superseded mailparser override.

Preserve Clack cancellation inference with its precise sentinel type and isolate the Vertex proxy fixture from ambient credentials. Timestamp and checksum audits, targeted consumers, native builds/tests, and independent review validate the refresh; required hosted CI remains the landing gate.

* fix(deps): preserve Clack cancellation types in exported prompts

Give styled configure prompts the exact upstream return types so plugin SDK declaration emission can name the new cancellation sentinel. Runtime behavior and generic option values are unchanged.

* fix(deps): preserve release tooling and Android test contracts

Regenerate Ruby lock metadata with pinned Bundler 2.6.9, grant Robolectric 4.17 its documented module access only in Android test JVMs, and keep the precise cancellation type without growing an over-cap source file.

Both previously failing Ruby lock guards, the line-cap and core type checks, all three configured Android test-task JVM arguments, and the actual Robolectric interceptor before/after probe pass. Independent review found no actionable P0/P1 issues.

* fix(deps): close Rustls advisory and align mock session clocks

Rustls 0.23.45 has now completed the seven-day cooldown; update only the shared crate pin and lock to the existing security-fixed desktop version.

Advance accepted mock Gateway writes on the synthetic fixture timeline and correlate permission tests with the actual mutation and refresh. This repairs a reproduced CI fixture race without changing production behavior or weakening assertions.

Validation: 36 Rust gateway-client tests including four TLS handshakes, 50 fixture tests, nine browser cases, scoped changed checks, and independent P0/P1 review passed.

* test(ui): keep external session updates on the committed timeline

* test: stabilize approval and desktop CI fixtures

* build(workboard): refresh assets after dependency rebase
2026-09-21 16:55:54 -07:00
Dallin Romney e8ea99af95 chore(qa): cover CLI status and health snapshots (#154613)
* test(qa): cover CLI status and health snapshots

* fix(qa): finalize retained gateway fixtures

* test(qa): audit CLI status health scenario
2026-09-21 15:52:53 -07:00
RoboClawandsteipete b86dc71207 refactor: remove compaction checkpoints (#154131)
* refactor: remove compaction checkpoints

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: d8185b46-c34b-46cb-a7bc-f4992867d3c9

* refactor: remove compaction checkpoints

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 80da0151-c7d1-4d40-80a6-e4a560bedf53

* refactor: remove compaction checkpoints

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: a7caf11d-e661-4a0f-8363-644cc007693a

* refactor: remove compaction checkpoints

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: e1a860df-02a1-4a99-a3f1-e9889eea119e

* test: repair checkpoint retirement CI fixtures

Route deferred QA tools through their declared dispatcher, preserve exact successful chronology receipts, and await rendered native accessibility values without weakening assertions.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test(codex): stabilize bounded rollout fixture

Reuse steipete's reviewed startup-scan synchronization repair from #154260 (89f04eec2c). Preserve preview, workspace, native-call-count and byte-limit assertions. Focused 12-case file passes on the composed candidate.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-20 19:47:11 -07:00