Files
OpenViking/examples/pi-coding-agent-extension/recall.ts
T
t0saki 674f5e6039 fix(retrieval): honor context tier ceilings and stop cooling unserved recalls (#3746)
* fix(retrieval): honor tier ceilings and stop cooling unserved recalls

Follow-up to #3534, from its post-merge review round.

- The abstract-to-overview substitute now applies only to categories whose
  stored abstract is the whole file body. A resource or skill whose abstract is
  missing (`processing_mode=vectors_only`) or over the per-entry cap read its
  body and returned an overview instead, which for a short file is the body
  almost verbatim — crossing the opt-in deepening boundary those categories are
  documented to have, and doing it even under an explicit `detail="abstract"`.
  They now degrade to a bare URI and their body is never read.
- A digest reporting `no_relevant` blanks `rendered`, so the client injects
  nothing, yet those URIs still entered the dedup ledger and were cooled for
  `dedup_turns` turns. That contradicted the ledger's own bare-URI grace rule
  and held memories back from the later turn they were relevant to.
- Flat retrieval reaches built-in memory types outside the four named ones
  (`cases`, `patterns`, `tools`, `trajectories`, skill-usage memories) and
  reported them as an undeclared `memories` category that no tier or penalty
  table covered, so other-peer hits skipped the score penalty and callers could
  not pin their tier. The catch-all is now a declared category with both; it
  stays out of `quotas`, whose buckets it would overlap. Skill-usage memories
  also stop being misread as the `skills` category.
- ZCode, OpenCode and pi own an OV session id but did not forward it, so their
  recalls silently ran without query expansion or cross-turn dedup.
- The context-request deadline covered only the server's 30s rewrite fuse, but
  the pipeline is serial: expansion, retrieval and budgeting all precede it.
  45s covers both fuses and the work between them.
- `plugin` config scope and the `/recall` successor example now match what the
  code actually does.

* fix(retrieval): make the context deadline and expansion opt-out reachable

Forwarding a session id turns on server-side query expansion, an LLM call with
its own 5s fuse, but neither the deadline that was supposed to cover it nor the
switch that turns it off reached the two harnesses this PR newly enabled it for.

- `contextRequestTimeoutMs()` now derives the deadline from the request body
  rather than from `cfg` plus a rewrite flag. The body is what states which
  server stages will run: a session takes the expansion fuse, `rewrite` takes
  the digest fuse, and a bare retrieval takes neither and keeps the caller's own
  budget. Reading `cfg` alone could not tell those apart.
- OpenCode pinned `timeoutMs: 5000` after spreading the helper's options and pi
  ignored them entirely, so the helper's deadline was dead code in both. Their
  own budgets are now defaults rather than ceilings. OpenCode's 5s in particular
  was shorter than the expansion fuse it had just enabled, so a legal request
  would have been aborted client-side and dropped back to the path with neither
  dedup nor expansion.
- OpenCode and pi read `OPENVIKING_RECALL_QUERY_EXPANSION` (and
  `recallQueryExpansion` in their own config files) and set the `configured`
  flag the shared body builder requires, so the documented opt-out exists where
  the cost was introduced.
- The integration overview no longer implies every harness reads the same
  environment knobs, and describes the deadline as per-stage rather than
  rewrite-only.
2026-08-05 23:52:15 +08:00

97 lines
3.1 KiB
TypeScript

import type { OVClient } from "./client.js";
import type { OVConfig } from "./config.js";
import { buildRecallBlock } from "./shared/recall-core.mjs";
export interface RecallCache {
block: string | null;
promptText: string; // the query this cache is for
}
export class RecallManager {
private client: OVClient;
private config: OVConfig;
private cache: RecallCache = { block: null, promptText: "" };
private pendingPrompt = "";
// Read lazily: the session manager that owns this id is constructed after the
// recall manager, and the id only exists once a session has been opened.
private sessionId: () => string | null;
constructor(client: OVClient, config: OVConfig, sessionId: () => string | null = () => null) {
this.client = client;
this.config = config;
this.sessionId = sessionId;
}
queueSearch(userQuery: string): void {
this.pendingPrompt = userQuery;
}
async searchPending(): Promise<string | null> {
if (!this.pendingPrompt) return this.cache.block;
const userQuery = this.pendingPrompt;
this.pendingPrompt = "";
if (userQuery.trim().length < this.config.minQueryLength) {
this.cache = { block: null, promptText: userQuery };
return null;
}
const block = await buildRecallBlock(
// 10s is this extension's own budget for a bare retrieval; when the
// request also spends a server fuse the helper hands down a longer
// deadline, and ignoring it would abort a request still inside its fuse.
(path: string, init?: any, options?: any) =>
this.client.fetchJSON(path, init, options?.timeoutMs ?? 10000),
this.config as any,
userQuery,
{
actorPeerId: this.config.peerId,
// Passing the OV session id is what turns on server-side query
// expansion and the cross-turn dedup ledger.
sessionId: this.sessionId() ?? "",
},
);
this.cache = { block, promptText: userQuery };
return block;
}
// --- Injection ---
injectRecall(messages: any[]): any[] {
if (!this.cache.block) return messages;
// Find the user message (scan backwards)
for (let i = messages.length - 1; i >= 0; i--) {
const msg = messages[i];
if (msg.role === "user") {
// Idempotency check
const content = typeof msg.content === "string"
? msg.content
: Array.isArray(msg.content)
? msg.content.filter((b: any) => b.type === "text").map((b: any) => b.text).join("")
: "";
if (content.includes("<openviking-context")) break;
// Prepend block to user message
const block = this.cache.block;
if (typeof msg.content === "string") {
msg.content = block + "\n" + msg.content;
} else if (Array.isArray(msg.content)) {
const textBlocks = msg.content.filter((b: any) => b.type === "text");
if (textBlocks.length > 0) {
(textBlocks[0] as any).text = block + "\n" + (textBlocks[0] as any).text;
}
}
break;
}
}
return messages;
}
invalidate(): void {
this.cache = { block: null, promptText: "" };
this.pendingPrompt = "";
}
}