mirror of
https://github.com/volcengine/OpenViking.git
synced 2026-09-30 09:17:51 +08:00
* fix(retrieval): honor tier ceilings and stop cooling unserved recalls Follow-up to #3534, from its post-merge review round. - The abstract-to-overview substitute now applies only to categories whose stored abstract is the whole file body. A resource or skill whose abstract is missing (`processing_mode=vectors_only`) or over the per-entry cap read its body and returned an overview instead, which for a short file is the body almost verbatim — crossing the opt-in deepening boundary those categories are documented to have, and doing it even under an explicit `detail="abstract"`. They now degrade to a bare URI and their body is never read. - A digest reporting `no_relevant` blanks `rendered`, so the client injects nothing, yet those URIs still entered the dedup ledger and were cooled for `dedup_turns` turns. That contradicted the ledger's own bare-URI grace rule and held memories back from the later turn they were relevant to. - Flat retrieval reaches built-in memory types outside the four named ones (`cases`, `patterns`, `tools`, `trajectories`, skill-usage memories) and reported them as an undeclared `memories` category that no tier or penalty table covered, so other-peer hits skipped the score penalty and callers could not pin their tier. The catch-all is now a declared category with both; it stays out of `quotas`, whose buckets it would overlap. Skill-usage memories also stop being misread as the `skills` category. - ZCode, OpenCode and pi own an OV session id but did not forward it, so their recalls silently ran without query expansion or cross-turn dedup. - The context-request deadline covered only the server's 30s rewrite fuse, but the pipeline is serial: expansion, retrieval and budgeting all precede it. 45s covers both fuses and the work between them. - `plugin` config scope and the `/recall` successor example now match what the code actually does. * fix(retrieval): make the context deadline and expansion opt-out reachable Forwarding a session id turns on server-side query expansion, an LLM call with its own 5s fuse, but neither the deadline that was supposed to cover it nor the switch that turns it off reached the two harnesses this PR newly enabled it for. - `contextRequestTimeoutMs()` now derives the deadline from the request body rather than from `cfg` plus a rewrite flag. The body is what states which server stages will run: a session takes the expansion fuse, `rewrite` takes the digest fuse, and a bare retrieval takes neither and keeps the caller's own budget. Reading `cfg` alone could not tell those apart. - OpenCode pinned `timeoutMs: 5000` after spreading the helper's options and pi ignored them entirely, so the helper's deadline was dead code in both. Their own budgets are now defaults rather than ceilings. OpenCode's 5s in particular was shorter than the expansion fuse it had just enabled, so a legal request would have been aborted client-side and dropped back to the path with neither dedup nor expansion. - OpenCode and pi read `OPENVIKING_RECALL_QUERY_EXPANSION` (and `recallQueryExpansion` in their own config files) and set the `configured` flag the shared body builder requires, so the documented opt-out exists where the cost was introduced. - The integration overview no longer implies every harness reads the same environment knobs, and describes the deadline as per-stage rather than rewrite-only.
20 lines
319 B
Plaintext
20 lines
319 B
Plaintext
{
|
|
"url": "http://localhost:1933",
|
|
"api_key": null,
|
|
"root_api_key": null,
|
|
"gateway_token": null,
|
|
"account": null,
|
|
"user": null,
|
|
"timeout": 60.0,
|
|
"output": "table",
|
|
"echo_command": true,
|
|
"extra_headers": null,
|
|
|
|
"plugin": {
|
|
"recallCompress": "auto",
|
|
|
|
"claude_code": {},
|
|
"codex": {}
|
|
}
|
|
}
|