Release-check lanes that were also red on main asserted pre-change contracts:
- subagent settle wording now says "in this batch" (#158642)
- Responses request bodies are pre-encoded bytes (#159574); decode them
without consuming init instead of requiring string bodies
- direct experience-review calls must bind the session MCP scheduler that
Gateway startup binds in production (#159481)
- Docker lanes pin the plain PATH `codex` CLI for the Codex harness probe
- Telegram group-policy hot reload must name every array it replaces
Landed directly from #160746 (the native publisher needs a signed merge to
bring that branch current). Focused proof: four QA suites 200/200 in 81 s on
the merged tree; the two E2E files and credentialed live paths were not run.
* fix(qa): recognize settled subagent batches
Share requester-settle wake recognition between mock input classification and completion handling. Accept the current batch wording while retaining installed-candidate session wording and the existing provenance check.
The real catalog-only handoff spawned and completed a child on the baseline, but QA ignored its completion and timed out. Two focused regressions failed before this fix. The real handoff now passes, all 70 owner/sibling tests pass, and check-changed passes. Standalone changed-test wall times: input 3.21s; handoff 33.76s.
* fix(qa): admit canonical repository checkpoint commands
* fix(ui): stabilize Run Inspector evidence collection
Bind rendered evidence to the public run, execution, and selected decision
receipt identities. Add a shared collector that uses the component-owned
route model and a separate page, preserving the caller's Chat surface and
unsent draft through collection and later reload.
Cover exact run/execution selection, receipt cursor reload, back navigation,
wrong requested identity, missing receipt, and page cleanup in the existing
mock-Gateway Chromium harness. Document the collector and per-tab auth setup.
The regression failed on baseline before product edits because the rendered
run identity attribute was absent. Terminal R's optional assistant transcript
idempotency key is a distinct producer gap; this does not manufacture that
key, relabel historical cells, or change product appearance.
* fix(telegram): confine QA runtime and preserve readiness evidence
Run standard-library drivers through existing Python without UV inline-script virtual environments. Confine private runtime state, drain readiness pipes, retain structural diagnostics, and require verified teardown receipts before releasing recovery state. Preserve doctor, recovery, group, and published-upgrade callers.
Validation: native sandbox regression failed before the fix; 140 harness tests pass, and 72 final focused owner/sibling tests pass in 5.73s. Build exits 0. Broad changed checks stop on two unchanged TS2459 package-update test errors. Focused lint matches all 135 baseline findings with zero additions. No live credentials or Telegram sends.
* fix(telegram): avoid native scenario barrier watchers
* fix(qa): share production publication guard contract
* fix(qa): require named message availability in mock provider
A catalog dispatcher does not advertise every delivery tool. Require the
exact message declaration and a usable invocation surface, including trusted
system/developer instruction carriers. Finish with the recovered child result
when message is absent.
Prove the absent-message regression through the mock HTTP provider and retain
catalog delivery, namespace, text-only fanout, and subagent handoff coverage.
* fix(qa): pass checkout roots to script scenarios
Expand the selected repoRoot separately from each scenario outputDir in the
maintained test-file runner. Prove the argument and CWD contract through a
real subprocess with external artifacts and fresh passing producer evidence.
* fix(qa): distinguish tool declarations from instruction prose
* fix(qa): bind forked-context evidence to native receipts
* fix(qa): bind repeated-ingress MCP scheduler
* fix(qa): add owned provider continuation checkpoints
Let maintained mock scenarios hold one session continuation before response
bytes and observe its request cursor and tool-call identity. Reuse the
provider request log and scenario signal/shutdown lifecycle; replacement
requests proceed normally without touching Gateway decision state.
* fix(qa): keep harness proofs within their owners
* test(qa): exclude all registered runtime consumers
* fix(qa): preserve runtime inputs and enforce checkpoint launches
* test(qa): skip runtime proofs when tools are absent
* fix(qa): complete checkpoint launcher test admission
* test(qa): compile repeated ingress child before execution
#160308 deleted src/agents/worktrees/service.run-end-cleanup.test.ts, which the
managed-worktrees-workboard-lifecycle scenario still listed as a codeRef, so
extensions/qa-lab/src/scenario-catalog.test.ts failed on main. The surviving
run-end cleanup outcome coverage lives in service.test.ts (late claims, stale
lifecycle writes) and service.removal-safety.test.ts (dirty retention).
Refresh application, plugin, native, build and container dependencies through the fixed 2026-09-19T16:27:11Z cutoff. Migrate native TypeScript snapshot/printer APIs while preserving compilation and filesystem contracts; retain existing patches and compatibility holds. Document offline container-image preparation.
Include the verified compiler process-census and loading-clock fixture repairs and deterministic warm-history regression. Adopt the canonical production history fixes from #159924 and #159955.
Land under the maintainer's explicit approval to treat proven pre-existing CI failures as non-blocking and repair main afterward. CI36365552098 failed an unchanged Android Compose fixture's asynchronous catalog projection assertion (3518 passed,1 failed); the Android/Gradle tree matches its main parent byte-for-byte. Security and dependency reviews passed. The final rebase preserves reviewed source changes and regenerates only the intentional Node-image documentation fingerprint. See PR159401 for complete validation and the follow-up repair obligation.
* fix(slack): keep top-level turns quiet and finish progress without "Working"
Surface: Slack channel plugin progress. Requested by Peter Steinberger.
Default progress turns without a reply thread now leave only the final
answer, using the temporary hourglass reaction when typingReaction is
unset and honoring configured reactions or an explicit disable. Preserve
explicit progress presentations and native-card selection for threads.
Finish untitled Block Kit cards as Done or Failed and native summary rows
as Completed or Failed, preserving authored headlines. Remove duplicate
Working fallbacks and build both session links from the dispatched session
key with the configured session.mainKey.
Derive terminal notification and screen reader fallback text from the same
rendered blocks, including queued-turn finalization, instead of reusing
working-state text.
Document the defaults and protect delivery, reaction cleanup, terminal
titles, and session links at the Slack wire and renderer boundaries.
Keep wire-trace helpers in test support to meet the existing line cap.
* test(qa): cover quiet top-level Slack progress
* test(infra): drain migration fixture databases before cleanup
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(slack): resolve verified linked requester profiles
Release note: Slack users linked to an administrative profile can ask an
agent to assign a session to them from ordinary messages and app mentions
received through Socket Mode or signature-verified HTTP. Generic trusted
requester metadata now carries the canonical linked profile ID and display
label; Discord verified senders benefit from the same core path. The
sessions tool remains owner-only.
Mint transport assurance at native reception and preserve it through the
existing durable ingress kind. Keep relay and mixed-assurance batches
asserted. Preserve the exact native Slack sender ID so the host identity
handoff matches finalized SenderId instead of rejecting its lowercase form.
Resolve requester facts through the existing worker-backed identity owner.
Bind their private carrier to the admitted sender, account, and channel,
recheck context and live link authority at prompt use, and retain no new
stored identity fields. Consolidate shared message-source and slash input
types to keep existing large modules within their line-growth limits.
Validation: 199 tests passed across five serial single-file Vitest runs,
including real Socket/HTTP receiver preparation, relay replay, linked
admin/member/unlinked senders, unlink/relink, genuine-carrier transplant
rejection, and persisted sessions.assignOwner through the agent tool.
Regressions failed before their corresponding repairs. Independent Codex
review is scoped-clean through P2; git diff --check passed.
The once-authorized check-changed run stopped at line-growth violations.
Those were repaired and affected tests rerun; the one-run host limit left
that gate unrepeated and typecheck/lint unrun. No live deployment was tested.
Follow-ups: native slash commands still need host-bound requester context;
assignment for non-owner channel requesters needs a separate permission design.
* fix(channels): reuse requester facts and keep prompts stable
Forward the single prepared linked-identity result, including known absence,
through command-owner authorization instead of resolving it a second time.
Keep the existing live authority checks and owner-only sessions policy.
Release note: linked requester profiles and short assign-to-me guidance now
live in host-generated per-turn conversation info, keeping the system prompt
byte-stable across different or unlinked senders. Use the canonical account
default and preserve lowercase Slack allowFrom and stored pairing approvals
while retaining native sender ID case for identity handoff.
Correct the cross-boundary Slack test fixture to use the maintained public
artifact loader, retain the context builder overloads, return synchronous
avatar data, and await preparation after SDK socket events. No SDK export,
protocol, configuration, schema, or permission expansion is introduced.
Validation: single-file authority tests 40 passed; Slack authorization tests
57 passed, including real persisted lowercase pairing approvals and a wrong-
sender negative control. New read-count and prompt-cache regressions failed
before the repair. Affected fixture cases passed again after type/lint fixes.
Full check-changed passed core/extension and test typechecks, lint, all guards,
and its six Doctor-contract tests. Independent Codex review of the final
staged candidate is scoped-clean through P2; git diff --check passed.
* test(qa): cover verified Slack requester assignment
Prove signed Slack HTTP ingress carries a linked administrative requester
profile in per-turn user context and offers the owner-only sessions tool.
An unlinked sender receives neither fact. Seed only the personal-profile
prerequisite while the ephemeral Gateway is stopped; use public role/link RPCs.
Capture strict Crabline readiness before runtime API traffic, retain the
unaltered runtime recorder, and probe the same adapter at final capture.
The recorder regression fails before the repair and passes afterward.
The optional model-issued assignment diagnostic still fails at the separate
admitted operator-authority boundary, despite a real shared target existing.
Record this coordinator gap without broadening permissions or claiming full
assignment success. Native slash requester context remains a follow-up.
Validation: signed-ingress QA passed on Blacksmith Testbox; focused tests,
applicable lint/typecheck/guards, and independent review through P2 passed.
The one full changed-file run exposed stale installed fs-safe; restoring the
pinned version fixed core types. Remaining core-test failure is fixed on main
by e9af8ebcde. No changelog, schema, protocol, or SDK surface changes.
* test(qa): register requester profile fixture entrypoint
Declare the process-launched QA fixture as a Knip executable root so its dependency graph remains audited. The exact failing unused-file scan now passes for both production and the full tree; independent review is clean through P2.
* fix(gateway): act as the linked admin for channel turns with verified owner authority
Carry the prepared linked administrator through the existing admitted operator authority. Preserve live command-owner, role, model, and access-grant restrictions for every dispatch and active model execution. Retain the original revocation failure for shared consumers. Prove visible-session assignment and fail-closed lifecycle behavior without changing tool gating, schemas, protocol, or SDK exports.
* test(qa): prove linked Slack admins can assign agent-owned sessions
* test(gateway): preserve original operator revocation assertions
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(tool-search): run tool_search_code in the QuickJS sandbox
Tool Search code mode spawned a Node --permission child with a node:vm
guest. Under Bun it needed an installed Node, and without one an explicit
code config silently downgraded to structured tools mode.
Run the guest through the Code Mode executor contract with the bundled
quickjs executor on every runtime. openclaw.tools.search/describe/call use
the shared namespace bridge with lazy thenables; the host admits only those
three operations through ToolSearchRuntime. codeTimeoutMs still bounds the
whole invocation, now including executor preparation. A denied or disabled
code-mode-quickjs plugin fails explicitly with next-step guidance instead of
falling back.
Remove the child source, IPC types, stderr-tail handling, and the Node and
Electron capability probe.
* refactor(tool-search): retire the tool_search_code bridge
Keep structured Tool Search and generic Code Mode as the two large-catalog surfaces. Remove the superseded JavaScript bridge and its runtime, display, and QA paths.
Doctor migrates legacy code mode to tools and removes codeTimeoutMs while preserving activation. toolSearch: true now selects structured search; JavaScript orchestration uses Code Mode exec/wait.
* test(tool-search): cover runtime behavior through structured controls
Exercise retained catalog, policy, hook, cancellation, terminal, MCP and client behavior through structured controls and their runtime owner. Delete bridge-only sandbox and JavaScript envelope cases while preserving nested call-id compatibility.
* test(tool-search): drop retired code mode prompt case
* test: repair fixtures exposed by Tool Search retirement checks
Remove the remaining retired Tool Search mode row. Preserve the session reader owner through the media retention mock and remove an unreachable queued-only branch from the ACP controls/submission fixture.
* test(upgrade-survivor): seed retired Tool Search code mode config
Author the legacy mode and timeout through every supported representative baseline CLI recipe, then require structured search and timeout removal after candidate update and Doctor. Existing config validation proves the resulting effective config.
Include the diagnostics native assignment summary in frozen target staging; the required assertion suite exposed its missing import. Node recipe and assertion tests: 238 passed, 157.87s wall. Docker validation remains with the coordinator.
* refactor(tool-search): drop the retired code-mode recovery surface
* chore: shrink assertion baseline after Tool Search retirement
* test(tool-search): drop the unused Tool Search test API
* chore: drop the retired Tool Search test API assertion baseline
* fix(e2e): drop duplicate native assignment staging line
Main now stages native-assignment-summary.mjs for frozen upgrades itself; the branch copy from the Tool Search upgrade proof became a duplicate after merging.
* build(pr): list Tool Search migration in wrapper inventory
The scripts/pr wrapper loads the Doctor config migrations at runtime, so the new Tool Search retirement migration belongs in its extracted component inventory.
* test(e2e): ship the Tool Search recipe to prepared tooling workers
Prepared tooling workers copy only listed source-relative assets, while the
upgrade survivor config recipe reads every section file by name at import.
The new tools-tool-search.json was missing, so the Docker scheduler parent
signal test's runner died with ENOENT and its polling wait reported a generic
5 s timeout that looked like a flake.
List the asset, guard the recipe directory against the preserved list, and
make the scheduler readiness wait fail with the runner's stderr once it exits.
Remove Tasks and TaskFlow runtime, APIs, CLI, SDK surfaces and panels after the Cron, session, native execution and media completion ownership cutovers. Preserve stored rows and import provable legacy native assignments through Doctor; ambiguous ownership stays untouched with a warning.
Follows #158221, #158217, #158225, #158222, #158702 and #158776. Related: #156532. Task-specific public APIs retire immediately; retained responsibilities use their existing owners.
Maintainer-authorized administrative landing after full CI run 36312986498 attempt 2 passed on 274595e2, with subsequent actual conflicts reviewed and focused checks passing. Current PR CI preflight hits the 64 KiB changed-path metadata limit before tests (run 36335042695); its duplicate security-review status mirrors that planning failure. Review and scoped proof are recorded in the PR. Published 9.4 native import is proven; remaining native completion and 9.4 rollback witnesses are explicitly unproven.
Point the thread memory and personal reply scenarios at the existing SDK protocol owner after the private forwarding barrel was removed. Reproduced the missing-reference failure on clean main; all 56 catalog tests pass after the two metadata fixes. Formatting and independent P2 review pass.
## What Problem This Solves
Fixes: owners who ask their agent in chat to change a key, a config value, or one of their own skills get refused, sent to a dashboard, or told to file a Workshop proposal. Examples: "you can't post API keys here", switching the embeddings provider routed to the web-search wizard, only proposals for handwritten skills.
## User Impact
An owner can hand the agent an API key or token in chat, ask it to change config such as the embeddings provider, and have it edit skills they own. Session permission modes are unchanged: Full Access applies, restricted sessions ask.
- **Keys from chat.** Masked setup flows still keep keys out of model context and remain the default. If the user already pasted a key or token, the agent stores it in the shared secret store and points the config key at it with a `store` SecretRef instead of refusing. It never echoes the value back. The pasted message already reached the model provider and transcript; redaction covers later logs and output only, and the docs say so.
- **Existing store entries are never touched.** Each save inserts a new entry named after the config key plus a random suffix (`GATEWAY_REMOTE_TOKEN_9B139B5E231299BC`). Nothing is overwritten, revived, or deleted, and a new name can never match anything already pointing into the store, including a stale reference to a removed and purged entry. Replacing a key leaves its previous entry for `openclaw secrets store rm`. Rotating a key keeps its configured store provider alias, and the audit records the alias actually used.
- **Embeddings.** The agent now treats the memory embeddings provider, model, and key as `memory.search.*` config, not the web-search setup wizard.
- **Skills.** When the user asks, the agent edits skills they own directly: repository skill source, workspace `skills/`, project `.agents/skills/`, and configured extra skill directories. Bundled, ClawHub-installed, and plugin-provided skills are replaced by their owners' updates. For those, the agent says so and offers to capture the change as a Workshop skill.
- **Tone.** The "never request / paste credentials in chat" lines are gone from the `openclaw` tools and system-agent prompts. "Never echo secret values" stays, and so do factual pointers for flows that genuinely need a UI: channel sign-in, provider OAuth/accounts, and model onboarding.
No config option, schema, or protocol change. The Full Access permission-policy floor from the first revision moved to #158142 for its own security review.
## Why This Change Was Made
After #149870, approved config writes may target any path. What still blocked owners was model-facing text telling the agent to refuse credentials, plus the missing ability to store a chat-provided value anywhere but plaintext config.
`config_set_ref` gains an optional `secret` argument (read without trimming; only emptiness is checked). With it, the system agent:
1. registers the value for redaction when the proposal is built;
2. keeps the key's existing store provider alias when it has one;
3. sends one `secrets.writeForConfigRef` command to the SQLite state worker with the requester's live-authority guard. The host re-checks that guard at the worker's transaction and commit admission (`createSqliteWorkerWriteAdmission`), so a run stopped while the command is queued writes nothing. The transaction inserts a new row under a freshly minted `NAME_<16 random hex>`;
4. writes the ref through the existing config writer, which re-checks authority. If that write fails (before or after the writer commits), OpenClaw rereads the config and the error says the key was saved as `<NAME>` and whether the config key points at it. There is no automatic delete: another consumer may have linked the fresh entry, or the writer may have committed before failing;
5. the normal config reload picks up the new ref, since its id always changes.
Nothing new runs SQLite on the Gateway main thread.
<details>
<summary>Out of scope / follow-ups</summary>
- Found while proving this: in Full Access, after a delegated change applies, the next agent turn in the same chat fails with `SQLite database already belongs to another worker backend`. It reproduces on unmodified `origin/main` (`71bb516`) with a `logging.level` change followed by one more message. This PR does not fix it.
- Built-in provider sign-in and model onboarding stay handoffs; they own live verification of the active inference route.
- Other secret-store set/delete paths remain synchronous migration debt, as `worker-access.md` already records.
</details>
## Evidence
Real Telegram Test Server (Convex-leased userbot, fresh Gateway, QA mock provider, Full Access, tester is owner), first revision:
| | Screenshot |
|---|---|
| Token given in chat, applied with no approval prompt and no refusal (synthetic QA token) |  |
### Final effects at this head
qa-channel scenario `system-agent-owner-trust` passes through a real Gateway and state worker. The Gateway is seeded with an unrelated `GATEWAY_REMOTE_TOKEN` entry, then:
1. A command-allowed **non-owner** (`bob`) sends the key. The `openclaw` tool is owner-only.
2. The **owner** (`alice`, Full Access) sends it.
Captured step details (redacted by the Gateway; the store ref id prints as `__OPENCLAW_REDACTED__`):
```json
{
"nonOwnerEntryNames": ["GATEWAY_REMOTE_TOKEN"],
"storedRef": { "source": "store", "provider": "default", "id": "__OPENCLAW_REDACTED__" },
"storeEntryNames": ["GATEWAY_REMOTE_TOKEN", "GATEWAY_REMOTE_TOKEN_9B139B5E231299BC"]
}
```
- After the non-owner turn: only the seeded entry exists and `gateway.remote.token` is unset.
- After the owner turn: `gateway.remote.token` is a `store` SecretRef, the token is in its own minted entry (`GATEWAY_REMOTE_TOKEN_9B139B5E231299BC`), and the seeded entry's `updatedAt`/`updatedBy` are unchanged. No approval prompt was posted, and the token is absent from chat, config, and the store listing.
**Revoked request**, through the production worker (Node main thread, real broker, `writeSecretStoreEntryForConfigRef`). The requester's guard passes the caller's check, then reports the run stopped:
```text
seeded: [ 'GATEWAY_REMOTE_TOKEN (cli)' ]
revoked request rejected: requesting run is no longer active
after revoked request: [ 'GATEWAY_REMOTE_TOKEN (cli)' ]
owner request saved as: GATEWAY_REMOTE_TOKEN_F1657B971F691824
after owner request: [ 'GATEWAY_REMOTE_TOKEN (cli)', 'GATEWAY_REMOTE_TOKEN_F1657B971F691824 (openclaw)' ]
seeded value intact: true
```
Tests (each fails without the behavior it covers):
- production worker path (`secret-store-config-ref.worker.test.ts`, forked database-worker lane with the real broker): a chat secret gets its own minted entry beside a live `GATEWAY_REMOTE_TOKEN` without touching it; a requester revoked after the caller's check writes nothing;
- store kernel: a refusal at commit admission rolls the transaction back; each save mints a new `NAME_<hex>` and leaves the key's previous entry unchanged; a stale name whose entry was removed and purged still resolves to nothing after a chat save;
- operations: stored and referenced with no value in output or audit; authority gone before the store write writes nothing; a failed config write names the saved entry and leaves it in place; rotating a key keeps its configured store provider alias, and the audit records it;
- tool: proposes a store write without repeating the key, preserving leading and trailing whitespace.
Measured single-worker wall time per new or materially changed test file at this head (`node scripts/run-vitest.mjs run <file>`, local M-series; vitest Duration includes import and setup):
| File | Tests | Wall | Vitest duration |
|---|---:|---:|---:|
| `src/secrets/store/secret-store-config-ref.worker.test.ts` (new, database-worker lane) | 2 | 14 s | 2.05 s |
| `src/secrets/store/secret-store.test.ts` | 32 | 16 s | 13.37 s |
| `src/system-agent/operations.test.ts` | 43 | 18 s | 15.30 s |
| `src/agents/tools/system-agent-tool.test.ts` | 36 | 15 s | 12.89 s |
QA scenarios: `system-agent-owner-trust` (mock-openai) runs in about 27 s after build; `skill-owner-direct-edit-live` is live-frontier only and took about 3 min with `claude-cli/claude-sonnet-4-6`.
Wording pins for the removed lecture text were deleted. The focused store, worker, exclusivity, operations, tool, approval, and delegate suites pass. `node scripts/check-changed.mjs` passes every gate except core lint, which fails only on three files this PR does not touch (`server-chat-metadata-lifecycle.integration.test.ts`, `session-companion-ask.ts`, `app-sidebar-session-list-render.ts` over `max-lines` on the base); oxlint on the changed files is clean.
Security decision: a Full Access owner's pasted key goes to the Gateway-wide team store without a separate approval. Maintainer (@obviyus) accepted this in the PR conversation.
**Rotation with a second consumer**, through the production worker (Node main thread, real broker). A second consumer references the key's first entry before the next save lands; value fingerprints only:
```text
owner saves key #1 -> MODELS_PROVIDERS_OPENAI_API_KEY_6D96F6E92B59E935 (sha256:4a5c5a4aa8de)
second consumer now references MODELS_PROVIDERS_OPENAI_API_KEY_6D96F6E92B59E935 (e.g. linked while the next save is queued)
owner saves key #2 -> MODELS_PROVIDERS_OPENAI_API_KEY_4066F18ABAAE6903 (sha256:28bc4e3fe10d)
second consumer's entry MODELS_PROVIDERS_OPENAI_API_KEY_6D96F6E92B59E935 after rotation: sha256:4a5c5a4aa8de
unchanged: true
```
**Config write fails after the save, with a second consumer on the fresh entry**, through the production worker and the system-agent apply path (fingerprints only):
```text
owner result: Saved the secret as GATEWAY_REMOTE_TOKEN_6B68C031E571F81C, but could not point gateway.remote.token at it: config write failed after commit (rollbackStatus: not-restored). Retry, or remove the entry with `openclaw secrets store rm GATEWAY_REMOTE_TOKEN_6B68C031E571F81C`.
config gateway.remote.token: null
second consumer's entry GATEWAY_REMOTE_TOKEN_6B68C031E571F81C: sha256:09ae5b4fd36b
second consumer keeps the credential: true
```
**Stale reference to a removed and purged entry**, production worker for purge and save:
```text
purged rows: 1
stale ref GATEWAY_REMOTE_TOKEN after purge: SECRET_STORE_NOT_FOUND
chat save -> GATEWAY_REMOTE_TOKEN_8F698B5396990ADD resolves (value hidden)
stale ref GATEWAY_REMOTE_TOKEN after chat save: SECRET_STORE_NOT_FOUND
```
**Owned-skill edit with a live model.** New scenario `skill-owner-direct-edit-live` (live-frontier; run with `claude-cli/claude-sonnet-4-6`, subscription auth) passes at this head. It seeds workspace skill `qa-owner-greeting` replying `OWNER-GREETING-V1`, and the owner asks in plain words: "Please change my qa-owner-greeting skill so it replies OWNER-GREETING-V2 instead of OWNER-GREETING-V1." Captured result:
Skill file after the turn:
```markdown
---
name: qa-owner-greeting
description: Greets the owner with a fixed marker
---
When the user asks for the owner greeting, reply with exactly: OWNER-GREETING-V2
```
Agent reply: "Let me find the skill file. Done. The `qa-owner-greeting` skill now replies `OWNER-GREETING-V2` instead of `OWNER-GREETING-V1`."
The model edited the skill file in place and confirmed it; it did not refuse or file a Workshop proposal.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* refactor(cron): read run history through its worker
* refactor(cron): clean up retired history type exports
* test(cron): finish history type and ratchet cleanup
* test(cron): prove published upgrade history reads
* test(cron): retain canonical failover type contract
* ci: refresh Cron checks after main import repair
Re-evaluate the unchanged reader against main after 7c4866c73b repaired the Team Reports test import. Production and upgrade-harness source remain unchanged.
* fix(cron): complete wrapper cutover and upgrade diagnostics
* test(cron): complete source references and delivery-free upgrade seed
* test(codex): keep catalog progress on one fake clock
* test(cron): align invalid run id assertion with owner
* ci: refresh Cron checks after async task main repair
* ci: refresh Cron checks after main QA reference repair
* ci: refresh Cron checks after Linux fixture repair
* ci: refresh Cron checks after workspace quota repair
* ci: refresh Cron checks after main progress type repair
* refactor: share filesystem admission, walks, and cleanup with fs-safe
* style: satisfy lint in fs-safe adoption changes
* test(daemon): map fs-safe reads into launchd fixtures
Keep the real fs-safe reader and its descriptor admission on the fixture backing files. Load fixture mocks before filesystem consumers and move shared state out of the oversized launchd suite.
* test(gateway): assert missing transferred manifest by errno
* fix(qa): drop retired Telegram formatter reference
The formatter refactor in #158164 folded format-render.ts into format.ts. Keep the existing format.ts code reference and remove the obsolete path so the scenario catalog reference check follows the current owner.
* fix(qa): preserve exact worktree cleanup identity
Keep the original bigint device and inode comparison when replacing the local forwarding helper. The general fs-safe matcher tolerates unknown Windows identity fields and cannot authorize destructive cleanup. Four real-directory receipt regressions fail the previous candidate; 35 cleanup/runtime cases and changed checks pass, with clean independent review.
## What Problem This Solves
Fixes: when an agent asks OpenClaw to change config, restart, or manage plugins/channels/agents, the approval never reaches chats without a native approval card, and `/approve <id>` rejects it. The agent's tool call then blocks until the 10-minute expiry, and users are sent to the Control UI or a terminal to approve a change they asked for in chat.
## User Impact
User impact: every approval an agent needs can be completed in the chat where it was requested.
- Native approval cards (Telegram, Discord, Slack, WhatsApp, Signal, Matrix, Teams, iMessage, Google Chat) keep owning their chats.
- Every other requesting chat now gets the change summary with **Allow once** / **Deny** buttons where the channel renders them, plus a `/approve <id> allow-once|deny` line.
- `/approve` now resolves OpenClaw change approvals, not only exec and plugin approvals.
No config options, protocol, or schema changes.
## Why This Change Was Made
Delegated OpenClaw changes ("system-agent" approvals) already had native cards (#134670). Two gaps remained outside those cards:
- **Request delivery.** The request was created with delivery turned off, so the shared approval forwarder never posted a fallback. The forwarder now has a system-agent strategy that always targets the requesting chat and is suppressed whenever that chat's native card handles the request. It uses the same typed-button payload and resolved/expired messages as exec and plugin approvals. The system-agent owner publishes the outcome of a decision (applied, or denied) once; approval publication only adds expiry and cancellation, so a denial produces one chat update. The fallback answers only the live messaging chat that made the request: terminal and Webchat requests never fall back to the session's saved chat, and expiry is reported once, from the Gateway's recorded expiry rather than a local chat timer, so a change approved just before the deadline and applied after it reports its applied outcome instead of a false expiry.
- **`/approve`.** The command probed only exec and plugin approvals. It now also resolves system-agent approvals through the canonical `approval.resolve`, after confirming the id is in `openclaw.approval.list`. Canonical resolution records a kind mismatch as a deny, so `/approve` confirms the owner first instead of probing. On channels with their own approver settings (Telegram, Discord, Slack, …), the Gateway checks the sender as the reviewer, the same check as native buttons. Everywhere else only a configured owner (`commands.ownerAllowFrom`) can approve an OpenClaw change: `/approve` sends the sender as the reviewer, and the Gateway checks owner custody against the current config inside the approval store's final decision guard. An ordinary command-authorized sender is refused, and an owner removed after sending `/approve` cannot complete the decision.
Unchanged: approval authority stays bound to the requesting run; free-text "yes" never approves; Full Access still auto-applies; Control UI and the apps can still decide.
Agent-facing text now says what happens: for runs from messaging channels, the `openclaw` tool description says the change waits for approval in this chat (buttons or `/approve`). Webchat and terminal runs, which the chat fallback cannot reach, are told to approve in the Control UI or OpenClaw apps. The gateway-only prompt line points to `openclaw` and `/restart` instead of "ask human".
## Evidence
Real Telegram (Test Server userbot, Convex-leased credentials, fresh Gateway, mock provider), restricted `exec` mode, agent asks OpenClaw to `set logging.level "info"`:
| Run | Setup | What the user saw | Tap | Result |
|---|---|---|---|---|
| A | tester is owner (native cards on) | 🔒 native approval card, no `/approve` text | **Allow Once** | `answerCallbackQuery` + 2× `editMessageText`: "approved. Applying" → "approved and applied"; final reply delivered |
| C | tester is owner, `channels.telegram.execApprovals.enabled: false` | fallback message with change summary, `/approve <id> allow-once\|deny`, and **Allow Once** button | **Allow Once** | message edited to "approved. Applying…", then "approved and applied" posted; Gateway config on disk has `logging.level: "info"`; final reply delivered |
Telegram Web screenshots from the same leased test account (cropped to the conversation):
| | Pending | After **Allow Once** |
|---|---|---|
| Native card (A) |  |  |
| Fallback message (C) |  |  |
With no owner configured, the owner-only `openclaw` tool isn't exposed, so no change is proposed (config unchanged). That's the existing design; DM pairing sets the first owner.
qa-channel (no native approval cards), new scenario `system-agent-chat-approval`: the agent's delegated change posts the approval in the requesting conversation. A command-authorized non-owner (`commands.allowFrom` includes them, `ownerAllowFrom` does not) sends `/approve <id> allow-once` and gets "❌ Only the owner can approve OpenClaw changes in this chat."; `logging.level` is unchanged and the approval stays pending. The owner's `/approve <id> allow-once` then resolves it; `logging.level` becomes `info`; exactly one final reply. On `origin/main` the same scenario times out waiting for the approval in the chat.
Tests: `/approve` (resolves a pending OpenClaw change canonically; never submits a canonical decision for an id owned by another kind; on a channel without approver settings, a non-owner and a revoked owner never reach the canonical decision), channel custody (without approver settings only a configured owner holds custody, and only for OpenClaw changes), `approval.resolve` (custody revoked between the request and the final write leaves the approval pending), approval publication (allowed/denied changes leave the chat outcome to the system-agent owner; expired/cancelled are published once), forwarder (requesting chat gets the prompt and outcome without `approvals.*` config; a running native card suppresses it; terminal and Webchat requests never reach the saved session chat; a change applied after the deadline reports its outcome with no false expiry; the recorded expiry is reported once), plus existing gateway approval and system-agent suites. The new owner-custody and single-publisher tests fail with the fix reverted. Shared approval and forwarder fixtures moved to sibling `*.test-support.ts` modules so the test files stay under the line cap. Wall time for the touched suites: `commands-approve` + `approval-publication` + `exec-approval-forwarder` + `system-agent-approval` + `approval` ran in 59.5s across 3 Vitest shards. `pnpm tsgo:core` clean.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* chore(deps): refresh dependencies with a seven-day cutoff
* fix(deps): preserve Teams and jsdom integration contracts
Use the Teams SDK public token and processing APIs while keeping SSO sender
checks ahead of native token operations. Remove obsolete ambient declarations
and route workarounds, and cover the SDK routing with real processing tests.
Adapt the test environment to jsdom private-field bindings, preserve file bytes
and registry cleanup, and preload it through native Node and Bun workers.
* fix(test): preserve jsdom window and fixture contracts
* fix(ci): keep typecheck cache reuse within matching inputs
* feat: show ready cloud worker pool in settings
* test: complete cloud worker cleanup fixtures
* test: wait for sidebar before plugin registration check
* test: run composed gateway scenarios in source lanes
* fix(gateway): opt in to live-authorized pool details
* test: document native TypeScript shutdown noise workaround
After mock request ownership moved to transport affinity, the mutating-tool
compaction scenario still searched prompt text for the session identifier.
The provider recorded the real context overflow, but those stale selectors
found no owning requests. A copied foreign identifier in prompt content
could also match the wrong session.
Select overflow, write, and continuation requests using their existing
request.sessionId field. Exercise the actual catalog expressions with the
owner identifier absent from prompt text and a foreign identifier copied
into it. Preserve the mutation, pruning, ordering, count, and failure checks.
No shared helper or production session contract changes.
The real scenario failed before the selector repair and passed all three
after samples: request size fell from 333954 to 119338 bytes, with exactly
one logical write, one authenticated wire success, and one compaction.
The changed catalog test passed 20 pressure runs at eight workers on two
CPUs plus a CPU contender. All 58 catalog consumers passed across the
four canonical CI shards (312 files, 4485 cases); the affected 78-file,
1062-case shard passed three times. Literal one-worker unit cost: 2.794s.
A fresh real scenario replay on main, including the guarded session
observer, passed in 21.309s. Proof runs: 35954197301 and 35957749551.
The changed checker completed all type graphs, guards and dead-code scans.
Remaining global lint was stopped in favor of scoped type-aware lint with
a rejecting negative canary; that substitute, import-cycle checks,
formatting, whitespace checks and independent P2 review passed.
Removing sessionId from the Runtime prompt intentionally stabilized the
cached prefix, but the QA mock still parsed it. Requests then shared
anonymous scenario state and terminal subagent settlement lost its
requester identity.
Observe full harness-generated ids through the existing before-run hook.
Resolve existing transport affinity by exact id, otherwise a unique
64-character prefix; reject unknown, ambiguous, or missing scoped ids.
Keep scenario state per session and expose canonical identity to the
subagent-completion scenario. Preserve affinity through QA proxies and
omit HTTP-only compatibility from the native harness catalog projection.
Update the settlement fixture for the published sessions.list contract
and align the package fixture with candidate-declared schema versions
and the existing explicit immediate-drain option.
The mock released child results when the parent's HTTP response was sent,
before the parent execution owner closed. Timestamp-prefixed all-settled
inputs also missed the fixture matcher and produced a generic response.
Correlate pending children with the exact runtime parent and use the existing
scenario wait loops to verify sessions.list reports that parent done, inactive
and not aborted. Keep the existing budgets and every direct-delivery, exact-send,
privacy and restart assertion. Recognize timestamped settlement inputs without
accepting quoted history. Migrate all mock-server and scenario callers together.
The six changed standalone test files each pass 20 full runs, with 91 focused
Gateway/QA cases and 43 HTTP sibling cases passing. Linux original profile 4
passes three times (27/27 scenarios, zero skipped), plus native Telegram and
QA-channel/empty flows. Actions proof: 35846151730. Single-worker file walls:
parser 2.47s, gate 2.48s, handoff 11.67s, routing 15.33s, surface 24.26s.
Types, scoped type-aware lint with a negative canary, formatting and P2 review
pass. Full lint declaration preparation hit ancestor-install isolation and used
the requested scoped substitute. Installed-package upgrade/rollback was not run
because no candidate tarball was configured; its migrated callbacks typecheck.
* refactor: retire pre-June import and verification compatibility
Remove pre-June task, flow, and plugin-state sidecar imports, obsolete
runtime chunks, package/installer validation exceptions, the old MCP
attachment fallback, and the April self-upgrade lane with its orphan helpers.
Leave retired data files untouched and document migration through 2026.6.1.
Preserve June-and-later contracts and September delivery recovery receipts.
Refs #156190
* docs: route legacy upgrades through 2026.9.5
* test: await Telegram fixture lifecycle events
Replace the setup stopwatch with the actual stop event or terminal run outcome. Keep cancellation assertions and outer execution bounds, and prove early terminal outcomes fail promptly.
* test(qa): repair forced restart and message inspection harnesses
* test(e2e): qualify installed execution identity persistence
Extend the packed npm onboarding owner with audit opt-in, one deterministic local turn, installed CLI inspection before and after Gateway restart, and isolated-state/privacy assertions.
Qualification from 267bb9d2: new shell-flow proof fails before the harness change at the missing opt-in. All 53 focused support tests pass; final changed shell/SQLite selection took 26.94s with one worker. Changed checks and P0-P2 autoreview pass; the full export scan was reused after brace-only lint fixes. Packed Docker proof remains a remote handoff.
* test(qa): qualify live execution and channel identity
* test(audit): qualify execution identity lifecycle gaps
* test(qa): align suppression audit with required replies
* test(qa): qualify Telegram participant identity
* test(qa): provision Telegram identity fixtures
* test(qa): qualify private-production Telegram identity
* test(e2e): preserve installed execution identity proof
* test(qa): await post-delivery memory maintenance
* fix(qa): repair qualification CI boundaries
## Summary
Unblocks the 2026.9.6 Full Release Validation (Release Checks run 35743792326, jobs 106800921622 and 106800921540): `compaction-retry-mutating-tool` failed identically in the parity (candidate) and runtime-pair (core) lanes with `Code Mode terminal continuation did not report successful completion`.
## Root cause
The scenario picked the expected terminal-continuation shape by provider variant: `openai` meant the Codex-native `Script completed\n...` text and `anthropic` meant the guest JSON `{ "status": "completed", ... }` result. The `Script completed` text only exists in the Codex harness (`extensions/codex`); the OpenClaw runtime's Code Mode always returns the guest JSON regardless of model.
Until #155614 an absent `tools.codeMode` defaulted to off, so this mock-openai lane performed the write through the direct `write` tool and the `writeWireToolName !== 'exec'` guard short-circuited the assertion. #155614 made the absent setting behave as `"auto"`, so `openai/gpt-5.5` now routes the write through guest `exec` and the never-satisfiable native branch fired. The product behaved correctly: the local repro shows the terminal continuation carrying `{"status":"completed", ..., "value":{"changed":true,"created":true,...}}`, the file with the exact expected content, and one compaction.
## Fix
- Mock OpenAI records the resolved Code Mode exec surface (`native` = freeform custom tool, `guest` = `code`-schema tool) on every request snapshot (`codeModeExecSurface`).
- The scenario keys the terminal-evidence assertion on that surface instead of the provider variant, still failing closed when neither surface is present. This is strictly stronger for the OpenClaw runtime: OpenAI-model guest runs are now actually verified for `status === "completed"` instead of being skipped.
- Catalog test pins updated to the new discriminator.
## Proof
- `node scripts/run-node.mjs qa suite --provider-mode mock-openai --parity-pack agentic --concurrency 1 --model openai/gpt-5.5 --alt-model openai/gpt-5.6-luna-alt --scenario compaction-retry-mutating-tool` on origin/main: fail (reproduced CI). With this change: pass, details `wireTool=exec ... wireSuccesses=1 compactions=1`.
- `scenario-catalog-compaction`, `scenario-catalog`, `mock-openai/server`, `agentic-parity-report` tests: 417 passed.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(qa): bind restart checkpoints to pending waits
* fix(sessions): preserve restart recovery safety guard
* fix(qa): honor current restart tool declarations
* fix(qa): honor unsafe restart tool declarations
* fix(qa): interrupt restart checkpoints before they drain
* fix(recovery): reject restart status before final capture
* test: align hot reload publication expectations
Reuse the exact three-line prerequisite from steipete PR155428 at 172e0aec90f44445ff24a4f1ae95af20993597c2. Full-file postimage and unchanged production contract match the donor; its four failing-case and six owner-supersession controls are reused. Supplemental API-direct full-file review is P0-P2 scoped-clean. No product behavior changes or assertion weakening.
* fix(models): retire catalog resources before closing work scope
Keep request cleanup admitted after credential work settles; drain the owner after discovery registry retirement. Preserve the existing cancellation barrier and normalize the equivalent Discord fixture to landed main.
* fix(qa): sanitize Slack failure causes without lint suppressions
Repair the production lint-suppression inventory failure from three
unlisted Slack QA directives. The existing failure sanitizer now creates
safe errors containing only validated Slack codes/scopes, preserving the
messages while dropping raw SDK headers and nested causes. Reuse the
existing Slack sender and retain exact channel receipt validation.
Validation: original lint inventory test reproduced the failure; repaired
inventory 4/4 and Slack owner suite 17/17 pass (28.33s combined wall with
one worker; Slack suite 9.14s total, about 0.7s test execution). Full
check:changed and independent P0-P2 review pass. No lint baseline changes.
(cherry picked from commit 44559eb315)
* test(qa): reconcile Slack runtime fixture with native send helper
Preserve the real Slack send helper in the adapter partial mock so deadline, cancellation, receipt and cleanup coverage executes native requests. Remove only the three obsolete inventory rows after landed sanitizer commit 44559eb315 removed their directives.
---------
Co-authored-by: Jason (Json) <263060202+fuller-stack-dev@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Related: #153456
## What Problem This Solves
Agents need reusable Discord and Slack E2E workflows with the same Convex-login setup as the Telegram userbot skill, rather than separate secret setup and ad hoc channel probes.
## User Impact
Developer-only: repository skills provide same-lease readiness, reusable YAML flows, native message/file/thread actions, correlated Gateway replies, and explicit evidence/cleanup instructions. Telegram reuses the extracted broker-discovery owner. Curated QA defaults and production channel behavior stay unchanged.
The dedicated Slack QA Driver now has the operator-approved reaction/file read/write scopes and was reinstalled. Its existing pooled token remains valid; no workspace-admin privileges, user-token scopes, or production-app changes were needed. The setup manifest includes those scopes.
## Why This Change Was Made
The skills extend the existing QA Lab transport, Gateway, scenario, and credential-lease owners instead of adding another runner. Opt-in `--doctor` and repeatable `--scenario-file` use Convex CI leases and deterministic `mock-openai` by default, while preserving explicit overrides.
Native-write flows do not retry ambiguous writes. Cleanup retains lease authority through Gateway shutdown, captures final native receipts before temporary state removal, and removes only owned fixtures. Native and cleanup Slack requests have bounded settlement with no write replay. Unanswered Gateway mutations remain explicit uncertainty and preserve runtime evidence instead of disappearing. Discord recording remains active through shutdown; threads are archived rather than deleted when the lease lacks thread-management permission.
Native API receipts are not model-tool or visual proof. Human slash commands, component clicks, modals, ephemeral replies, and rendering retain explicit client-testing requirements. Broader testing capabilities are deferred to follow-up PRs.
## Evidence
- Real Discord full native lifecycle passed on `ea1b0536204`: messages/replies/pagination/edit, reaction add/remove, public thread, attachment upload/delete, a fresh correlated SUT reply, continuous recording, and successful owned cleanup.
- Real Slack full native lifecycle passed on `db88f0ac1cc`: all five steps passed, including quiet ingress, threaded Gateway reply, stored edit/delete and reaction add/remove verification, and exact file upload/readback. Cleanup reported zero remaining owned messages/files/reactions and zero pending, uncertain, or failed operations.
- Both native runs used real transports, a temporary Gateway, `mock-openai`, and Convex CLI discovery with both broker environment variables unset. Telegram/shared discovery regressions: 30 passed.
- Earlier focused QA set: 301 passed; subsequent affected-consumer set: 298 passed (overlapping, not additive). Extension typecheck, changed-TypeScript type-aware lint, source ratchets, docs checks, and whitespace checks passed during preparation.
- Landing reconciled one pinned main snapshot, `eac221ae78bd`, preserving both lifecycle guards. After restoring its frozen dependency versions, the merged Gateway lifecycle, Slack ownership/adapter, and scenario deadline suites passed: 4 files, 52 tests. Earlier native artifacts are retained as pre-reconciliation proof, not relabeled as executions of the merge head.
- Repaired-head Slack lifecycle passed on `62ed1326be41`: all five steps, reaction/file readback, final receipt capture, and clean owned teardown. This rerun covers the bounded-request and uncertain-capture repairs, on the reconciled base.
- Both teardown regression controls failed on the pre-repair implementations (missing HTTP deadline; unanswered mutations dropped). The repaired owner/consumer set passed all 73 tests across six files. Targeted type-aware lint, extension typecheck, docs MDX and whitespace checks also passed.
- Hosted CI will not be awaited, per the operator's explicit landing instruction; no full CI success is claimed.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* chore(deps): refresh dependencies with seven-day cutoff
Advance eligible runtime, native, release, and development dependencies published by 2026-09-14T07:00:00Z. Preserve compatibility holds and existing reviewed newer pins. Synchronize release integrity checks and scoped overrides; remove the superseded mailparser override.
Preserve Clack cancellation inference with its precise sentinel type and isolate the Vertex proxy fixture from ambient credentials. Timestamp and checksum audits, targeted consumers, native builds/tests, and independent review validate the refresh; required hosted CI remains the landing gate.
* fix(deps): preserve Clack cancellation types in exported prompts
Give styled configure prompts the exact upstream return types so plugin SDK declaration emission can name the new cancellation sentinel. Runtime behavior and generic option values are unchanged.
* fix(deps): preserve release tooling and Android test contracts
Regenerate Ruby lock metadata with pinned Bundler 2.6.9, grant Robolectric 4.17 its documented module access only in Android test JVMs, and keep the precise cancellation type without growing an over-cap source file.
Both previously failing Ruby lock guards, the line-cap and core type checks, all three configured Android test-task JVM arguments, and the actual Robolectric interceptor before/after probe pass. Independent review found no actionable P0/P1 issues.
* fix(deps): close Rustls advisory and align mock session clocks
Rustls 0.23.45 has now completed the seven-day cooldown; update only the shared crate pin and lock to the existing security-fixed desktop version.
Advance accepted mock Gateway writes on the synthetic fixture timeline and correlate permission tests with the actual mutation and refresh. This repairs a reproduced CI fixture race without changing production behavior or weakening assertions.
Validation: 36 Rust gateway-client tests including four TLS handshakes, 50 fixture tests, nine browser cases, scoped changed checks, and independent P0/P1 review passed.
* test(ui): keep external session updates on the committed timeline
* test: stabilize approval and desktop CI fixtures
* build(workboard): refresh assets after dependency rebase