mirror of
https://github.com/openclaw/openclaw.git
synced 2026-09-28 05:54:09 +08:00
Closes #150575 ## What Problem This Solves Telegram mention-gated groups lose useful discussion outside the recent context window, and the agent cannot retrieve that older discussion from the existing message cache. ## User Impact The agent receives a bounded recent window (existing `historyLimit`, default 50), then can use `message(action="read")` to retrieve earlier permitted discussion instead of receiving thousands of messages in its prompt. Unmentioned messages do not start agent turns. Explicit replies, existing transcripts, direct messages, and `requireMention: false` activation remain supported. This reuses the existing SQLite table and primary-key index: no new table, column, index, database, schema-version bump, Doctor migration, or configuration key. Recorded group history survives restarts and remains readable after `/new` or `/reset`. `historyLimit: 0` disables automatic context, not explicit reads. **Retention and downgrade tradeoff:** group records no longer expire by message count, so disk usage grows with observed traffic. Only messages delivered to and permitted by OpenClaw are available; this is not Telegram server-side backfill. Older binaries retain their old eviction/quota behavior and must not write to expanded retained history. Restore a compatible pre-update backup before downgrading; this change intentionally does not add a database-version fence. ## Why This Change Was Made - Use one durable Telegram message owner for recent context, explicit reads, edits, and reply metadata; remove the separate in-memory group-history buffer. - Add generic retained namespaces, indexed key-range paging, and atomic promotion to the existing plugin-scoped store. Telegram is the first consumer; other channels can adopt these primitives without Telegram-specific core policy. - Keep reads scoped to the trusted current account, chat, and topic, with native message-ID cursors and permission/session/plugin authority rechecked after storage awaits. Embedded reply snapshots do not become authorized history merely because they are cached. - Move eligible legacy group rows atomically before bounded writes can evict them; bounded stores and ordinary DM caching retain their existing limits. Related prior work: #99143 and #121907. This implements the hybrid selected in #150575 rather than reopening the closed retrieval-only proposals #143665, #143892, or #147799. Other-channel fixture edits adapt to the shared async opener's option union; they do not enable retained history for those channels. ## Evidence ### Real Telegram Test Server, leased userbot Used the repository's Telegram E2E userbot driver with a deterministic mock model provider, real user sends, the normal Gateway, and the registered message tool: - 55 quiet group messages: no SUT event before the mention; the first provider request included exactly messages 5–54 (50 messages), not the early fact in message 0. - The registered `message read` action returned all 55 earlier messages, including the omitted fact. - After Gateway restart and session reset, explicit reads still recovered the earlier fact; old observations did not repopulate the reset automatic window. - A separate `historyLimit: 0` scenario included no automatic observation rows, while explicit reads worked before and after restart/reset. - Both completed runs removed credential scratch. The run-owned group was deleted by the driver. No operator Gateway was modified. ### Storage, policy, and regression proof - Focused cache/recovery/shared-dispatch proof: 25 tests passed. Covers first-DM legacy promotion across scopes, more than one promotion batch, promotion failure without eviction, explicit reply/media preservation with zero automatic history, numeric paging, scope isolation, authority revocation during awaits, restart/reset reads, and edit/media races. - All 26 selected Telegram/shared-dispatch test files passed after repairs, including bot entry points, quiet history, native commands, account/topic isolation, media and external replies, albums, send/edit/poll behavior, and action discovery. The final remaining batch passed 442 tests across 13 files. - 157 storage tests passed, including more than 50,000 retained records, reopening, cold bounded sibling writes, atomic promotion, and runtime capability revocation. The retained paging query uses the existing primary-key index. - Final production build passed in a separate physical checkout, including SDK declarations and verification of all 154 public SDK subpaths. Core, extension, extension-test, root-test, and core-test typecheck lanes passed. Type-aware lint for changed files, documentation MDX checks, line limits, assertion-safety ratchet, and whitespace checks passed. - Independent review identified two P2 edge cases (DM-first legacy eviction and zero-window topic recovery); both were fixed, with failing-before/passing-after regression proof. A supplemental review of the subsequent production fixes found no actionable P0–P2 issues. - Known baseline limitation: `plugin-sdk:surface:check` reports 137 deprecated `channel-message` exports against a budget of 136 on both the clean pinned baseline and candidate. No budget increase or suppression was added. ### ClawSweeper follow-up Both correctness findings on `e05ef2b93dc4` are repaired: - Automatic assembly uses the existing canonical session boundary rather than searching archived text for reset commands. It inspects at most `max(256, historyLimit)` raw records before topic and sender filtering. Sparse or filtered conversations can yield fewer automatic messages; explicit history pagination remains unchanged. Recovered topics carry the same canonical reset boundary. - Explicit reply traversal prefers canonical stored ancestors, then uses an embedded legacy snapshot only on a missing key. The existing depth/cycle and conversation visibility limits remain; lookup errors propagate, and fallback neither writes ancestor rows nor grants ordinary-history eligibility. Measured against the old head using the same 50,512-row SQLite fixture: | Automatic workload | Before: stored rows returned by range reads | After | | --- | ---: | ---: | | Default 50-message window, no reset markers | 50,768 | at most 256 | | `historyLimit: 0`, no reply | 50,512 | 0; no history-storage reads | | Current sender-policy filtering | 101,843 | at most 256 | | Canonical reset boundary after reopening | 50,768 | at most 256 | The sparse-topic regression also verifies that automatic selection stops at its physical window while explicit reads retrieve older matching records. The upgrade regression starts with a legacy parent `9` whose only copy of ancestor `8` is embedded, promotes and reopens the database, then recovers `9 → 8`. On the old head it returned only `9`. The raw parent payload remains unchanged; no standalone key for `8` or general-history eligibility is introduced. Compatibility is unchanged: these repairs change query/traversal behavior, not tables, columns, indexes, persisted payload versions, configuration keys, retention policy, or the documented backup-only downgrade restriction. Validation for this revision: 78 focused tests plus 187 sibling tests passed; extension source/test types, targeted type-aware lint, line-cap/assertion-safety checks, documentation parsing, whitespace checks, and the QA runtime build passed. Independent review of the production repair found no actionable P0–P2 issues. Fresh leased Telegram Test Server userbot runs repeated both default-50 and zero-window flows through the normal Gateway and registered read tool, using the deterministic mock provider. The default run supplied exactly messages 5–54 after 55 quiet sends; the older fact was absent from automatic context but present in explicit reads. Both runs recovered that fact after Gateway restart and `/new`, with no SUT activity before the mention. Both runner exits were zero, credential scratch was removed, and the run-owned group was deleted. ### Ready-for-landing preparation Integrated pinned main `f31305a6cdad` to resolve a real merge conflict. The conflict was an upstream SQLite helper relocation; the extracted database owner now imports that helper from its new canonical module. New upstream ClickClack fixtures were adapted to the async-store option union without forwarding retained options into bounded synchronous stores. On the integrated source, 323 storage/Telegram tests and 56 ClickClack tests passed, along with core, extension, and extension-test typechecks. The Telegram history implementation and its retention/upgrade contract are unchanged by this integration. The retained live userbot captures were taken on `faabc264ebd`; the integration adds the targeted storage, reaction-routing, rich-message, and fixture proof described here. Hosted CI on `3244f91f353` identified three integration guard failures. Removed the now-unused Telegram media-ID helper; updated the registered provider-owned read-gate contract to include the implemented `read` action; removed the obsolete source-text assertion requiring the old in-memory history facade while retaining the guards against deprecated low-level history helpers. No runtime behavior, retention policy, or permissions were weakened. The repaired plugin-shape, media, and channel-turn suites passed 174 tests, and both complete Knip dead-code passes succeeded. ### Accepted-send history-failure settlement A subsequent exact-head review identified a post-send settlement gap. Telegram could accept a native-quote answer, then the final history write could escape as an ordinary unsent failure. The prepared sender's existing acceptance ledger now supplies the reply receipt. Final history failures preserve accepted IDs and provider-observed placement through the existing typed partial-delivery outcome; failed cleanup preserves both failure causes. Finalized previews use the existing partial-finalization result rather than losing their accepted receipt. Storage errors remain errors, not successful retention. Discriminating regressions use the real shared dispatcher and actual Telegram delivery implementation with an accepted API boundary and a failing history writer. On `de0b40b1ef9`, the single-answer case made 2 API sends instead of 1; the multi-chunk and failed-send/failed-cleanup cases made 6 instead of 2. The repaired cases preserve the exact accepted prefix and receipt without fallback/resend. The old preview path threw instead of retaining its finalized receipt. All 223 delivery/receipt tests plus 75 native-command/reply-target/progress siblings passed; source/test types, targeted lint, complete dead-code checks, size/safety gates, and the runtime build passed. A fresh leased Telegram Test Server run injected a targeted SQLite write fault in runner-owned state only. One native-quoted answer arrived, its history row remained absent, the write error stayed visible in diagnostics, and no duplicate or fallback was observed. Fault triggers were removed before shutdown; the owned group and credential scratch were cleaned up. This live default durable-ingress lane also suppresses fallback on the old head, so it is not presented as the negative control: the discriminating before/after fallback and typed-receipt evidence is the shared-dispatch/API-boundary regression above. ### Upstream CI prerequisite integration Head `0ec9b6a3c6f` integrates main through `3fceb86048c4`, which contains the upstream repair for duplicate `promisify` declarations in three SQLite fixtures. The previous CI run tested the synthetic merge with `47c4fbcb20d2`; that base contains the same duplicate declarations, and its own CI also failed the corresponding type, lint, boundary, and test jobs. No check was waived. All 81 tests in the three repaired SQLite suites pass on the integrated candidate. ### Final readiness checks The refreshed [CI run 35394554718](https://github.com/openclaw/openclaw/actions/runs/35394554718) on `0ec9b6a3c6f5efec6829da077a32dcd7d69891f8` cleared the earlier SQLite syntax, type, lint, boundary, and Linux/Windows test failures. Only `QA Smoke CI (profile 4/4)` and its aggregate `openclaw/ci-gate` remain red. The remaining scenario failure is independently reproduced on pinned upstream main `3fceb86048c49eebc244ed68205baf7d3daa4fe5`, without this PR. The baseline used an isolated source archive, its own physical dependency installation, the private QA runtime build, and the same smoke-profile scenario with Crabline Telegram and the deterministic mock provider: ```text OPENCLAW_BUILD_PRIVATE_QA=1 pnpm build qaRuntime pnpm openclaw qa run --repo-root . --qa-profile smoke-ci \ --scenario subagent-completion-direct-fallback --output-dir .artifacts/baseline-qa Baseline: 0 passed, 1 failed, 0 skipped; exit 1 [private] timed out after 90000ms qa-terminal-private-first: status=completed, deliveryStatus=pending Private parent continuation absent; second private child never created. Public visible, silent, and fallback controls: completed and delivered. ``` The baseline and both PR CI runs exhibit the same private-completion failure after the scenario's Gateway restart. This classifies the remaining QA failure as an upstream baseline defect, not a retained-history regression. No retry, timeout increase, assertion weakening, CI waiver, or merge bypass was added. That prior CI gate remained red; landing requires evaluating the new head and resolving any remaining baseline failure, subject to the current exact-head landing gates. ### Aggressive reduction pass Cleanup head `60e4c4a404c` removes **1,988 net lines** relative to the prepared implementation: **421 production lines** and **1,567 test/support lines**. The main cleanup commit removes 3,100 lines and adds 1,112; a separate generated baseline update only lowers the store assertion allowance. At that checkpoint, the whole PR was **+1,392 net lines**, down from +3,380 (59% less net growth), not net negative. Removed the orphaned store health probe and its self-tests, the global store-test bridge, the unused cache `includeNode` path, duplicate storage error envelopes, redundant delivery bookkeeping, and overlapping persistence/history/delivery tests. Consolidated prompt fixtures without preserving their casts or manual temporary-directory cleanup. Unique permission/revocation, legacy promotion/ancestry, bounded scan, reset, quota, concurrency, and accepted-send regressions remain. No configuration, schema, retention, downgrade, permission, or delivery contract changed. Verification: - 889 tests across 33 selected store/Telegram/registered-dispatch files passed; the corrected typed prompt fixture additionally passed its 11 tests. - Core and extension source types, core/extension/root test types, targeted type-aware lint, line-cap/max-lines checks, exact shrink-only assertion ratchet, whitespace, and the isolated QA runtime build passed. - Independent P0–P2 review of the reduction found no actionable findings or lost unique critical coverage. - Fresh leased Telegram Test Server runs used real user messages, the normal Gateway, strict readiness, run-owned groups, and the registered message tool. After 55 quiet messages, the default run supplied exactly messages 5–54; the zero-window run supplied none. Explicit reads recovered the omitted fact initially, after Gateway restart, and after `/new`. Both reset automatic windows omitted the old observations. Neither run emitted SUT activity before the mention. Both exited zero, deleted their owned group, removed credential scratch, and stopped their listeners. The earlier private-subagent QA failure remains documented against its exact prior CI run above; it is not waived by this cleanup. Fresh hosted checks on the new head must be evaluated separately before landing. ### Final main integration Head `fc413c314e5` integrates pinned main `330f3c2f8352`, resolving the real merge conflicts without restoring deleted test scaffolding or competing history paths. Main's #152363 removes the aggregate plugin row quota, so this PR now preserves namespace-only bounded limits and removes obsolete aggregate retained-versus-bounded bookkeeping and fuse tests. Retained namespaces, raw atomic promotion, scoped reads, and backup-only downgrade recovery remain unchanged. The new upstream namespace-independence regression remains registered. Against this pinned main, the whole PR is **+1,236 net lines**: +766 production, +428 tests/support, and +42 docs/config. This remains net positive; the cleanup did not delete unique behavior or unrelated tests merely to reach a negative total. Integration proof: 281 selected store/Telegram/registered-dispatch tests and all four new upstream Crabbox authority cases passed. Source and test typechecks, targeted type-aware lint, MDX parsing, size/safety ratchets, and the isolated QA runtime build passed. A fresh independent P0–P2 review of the complete integrated PR returned no actionable findings across all three review passes. Both leased Telegram Test Server userbot flows were repeated on the final integrated source bytes: 55 quiet messages produced exactly 50 automatic messages by default and zero with `historyLimit: 0`; the registered read tool recovered the omitted fact initially, after restart, and after `/new`. Reset automatic windows stayed empty, and no SUT activity preceded the mention. Both runs exited zero, deleted their owned groups, removed credential scratch, and stopped their listeners. [Current-head CI run](https://github.com/openclaw/openclaw/actions/runs/35418671923) is attached; hosted results remain a separate landing gate. The author has authorized landing after exact-head review and CI gates. ### Landing gate repair The final candidate `1bef6da8a7b` makes the database opener private and routes fixture clearing through the existing transaction helper. This removes the unused production export reported by hosted `check-dependencies`; both complete Knip passes and 48 focused storage/Doctor/error tests now pass. No Telegram runtime logic changed, so the final integrated userbot proof above remains applicable. All four QA smoke profiles passed on `fc413c314e5`, superseding the earlier baseline private-completion timeout. [Final candidate CI](https://github.com/openclaw/openclaw/actions/runs/35419666431) is the required landing check. No CI waiver or merge bypass is requested. Landing is authorized; the previous author hold is released. Co-authored-by: Ayaan Zaidi <hi@obviy.us>
130 lines
5.1 KiB
TypeScript
130 lines
5.1 KiB
TypeScript
import { createHash } from "node:crypto";
|
|
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
|
import { createPluginRuntimeMock } from "../src/plugin-sdk/test-helpers/plugin-runtime-mock.js";
|
|
import type {
|
|
OpenAsyncKeyedStoreOptions,
|
|
OpenKeyedStoreOptions,
|
|
} from "../src/plugin-state/plugin-state-store.js";
|
|
import { createTestRegistry } from "../src/test-utils/channel-plugins.js";
|
|
import { withOpenClawTestState } from "../src/test-utils/openclaw-test-state.js";
|
|
|
|
beforeEach(() => vi.resetModules());
|
|
|
|
afterEach(async () => {
|
|
const { resetPluginRuntimeStateForTest } = await import("../src/plugins/runtime.js");
|
|
resetPluginRuntimeStateForTest();
|
|
vi.restoreAllMocks();
|
|
});
|
|
|
|
describe("iMessage persisted alias matching through the registered adapter", () => {
|
|
it.each(["modern", "legacy"] as const)(
|
|
"matches a cold SQLite current-message binding on the %s host path",
|
|
async (host) => {
|
|
await withOpenClawTestState({ label: `imessage-alias-${host}` }, async (state) => {
|
|
const { createPluginStateKeyedStore, createPluginStateSyncKeyedStore } =
|
|
await import("../src/plugin-state/plugin-state-store.js");
|
|
const entry = {
|
|
accountId: "work",
|
|
messageId: "00000000-0000-4000-8000-000000000042",
|
|
shortId: "7",
|
|
chatId: 42,
|
|
chatIdentifier: "person@example.test",
|
|
timestamp: Date.now(),
|
|
};
|
|
// Seed the persisted upgrade contract before loading the plugin's memory cache.
|
|
createPluginStateSyncKeyedStore("imessage", {
|
|
namespace: "imessage.reply-cache",
|
|
maxEntries: 2000,
|
|
env: state.env,
|
|
}).register(
|
|
createHash("sha256").update(entry.messageId).digest("hex").slice(0, 32),
|
|
entry,
|
|
{ ttlMs: 6 * 60 * 60 * 1000 },
|
|
);
|
|
createPluginStateSyncKeyedStore("imessage", {
|
|
namespace: "imessage.reply-cache-counter",
|
|
maxEntries: 1,
|
|
env: state.env,
|
|
}).register("short-id-counter", { counter: 7 });
|
|
|
|
const runtime = createPluginRuntimeMock({
|
|
state: {
|
|
resolveStateDir: () => state.stateDir,
|
|
openKeyedStore: <T>(options: OpenAsyncKeyedStoreOptions) =>
|
|
createPluginStateKeyedStore<T>("imessage", { ...options, env: state.env }),
|
|
openSyncKeyedStore: <T>(options: OpenKeyedStoreOptions) =>
|
|
createPluginStateSyncKeyedStore<T>("imessage", { ...options, env: state.env }),
|
|
},
|
|
});
|
|
const openSync = vi.spyOn(runtime.state, "openSyncKeyedStore");
|
|
const { imessageMessageActions, setIMessageRuntime } =
|
|
await import("../extensions/imessage/runtime-api.js");
|
|
const { imessagePlugin } = await import("../extensions/imessage/api.js");
|
|
setIMessageRuntime(runtime);
|
|
const delivered = {
|
|
content: [{ type: "text" as const, text: "reaction delivered" }],
|
|
details: { ok: true },
|
|
};
|
|
const handleAction = vi.fn(async () => delivered);
|
|
const { setActivePluginRegistry } = await import("../src/plugins/runtime.js");
|
|
setActivePluginRegistry(
|
|
createTestRegistry([
|
|
{
|
|
pluginId: "imessage",
|
|
source: "test",
|
|
origin: "bundled",
|
|
plugin: {
|
|
...imessagePlugin,
|
|
actions: { ...imessageMessageActions, handleAction },
|
|
},
|
|
},
|
|
]),
|
|
);
|
|
const matchParams = {
|
|
args: { chatId: 42, messageId: entry.messageId },
|
|
accountId: "work",
|
|
toolContext: {
|
|
currentChannelProvider: "imessage" as const,
|
|
currentChannelId: "person@example.test",
|
|
currentMessageId: entry.shortId,
|
|
},
|
|
};
|
|
|
|
if (host === "legacy") {
|
|
const match =
|
|
imessageMessageActions.messageActionTargetAliases?.react?.matchesCurrentConversation;
|
|
expect(match?.(matchParams)).toBe(true);
|
|
expect(match?.({ ...matchParams, args: { ...matchParams.args, chatId: 99 } })).toBe(
|
|
false,
|
|
);
|
|
expect(openSync).toHaveBeenCalled();
|
|
expect(handleAction).not.toHaveBeenCalled();
|
|
return;
|
|
}
|
|
|
|
const { dispatchChannelMessageAction } =
|
|
await import("../src/channels/plugins/message-action-dispatch.js");
|
|
const context = {
|
|
cfg: {},
|
|
channel: "imessage" as const,
|
|
action: "react" as const,
|
|
params: matchParams.args,
|
|
accountId: "work",
|
|
requesterAccountId: "work",
|
|
conversationReadOrigin: "delegated" as const,
|
|
toolContext: matchParams.toolContext,
|
|
};
|
|
await expect(dispatchChannelMessageAction(context)).resolves.toBe(delivered);
|
|
await expect(
|
|
dispatchChannelMessageAction({
|
|
...context,
|
|
params: { ...context.params, chatId: 99 },
|
|
}),
|
|
).rejects.toThrow("exact current conversation");
|
|
expect(handleAction).toHaveBeenCalledOnce();
|
|
expect(openSync).not.toHaveBeenCalled();
|
|
});
|
|
},
|
|
);
|
|
});
|