* feat: add freshness-aware parent aggregation
Defer wide-directory abstract/overview regeneration until the configured freshness threshold is reached while continuing changed-file semantic and vector processing.
Persist freshness metadata atomically, make parent bubbling L0-aware, preserve separate semantic/vector statuses, and keep explicit waits synchronous.
Rebuild every sampled summary on threshold refresh and always retry directory vectorization so stale sidecars or transient vector failures cannot be silently accepted.
Add focused coverage for freshness policy, pending-state consumption, sampled-summary refresh, vector retries, and parent bubbling.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive, and applied to memory
* feat: ov reindex support --recursive, and applied to memory
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Materialize session-aware context requests through SessionService before
loading the recall ledger, while retaining the messages.jsonl guard for
partially initialized sessions.
Add coverage for first-turn ledger writes, same-session message capture,
materialization failure recovery, and stateless requests with session
features disabled.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Treat messages.jsonl as the materialization boundary for session-aware recall,
repair partial session roots during the existing authoritative append path,
and preserve Claude capture cursors when writes never reach the server.
Also replay explicitly retryable storage conflicts across memory plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): honor tier ceilings and stop cooling unserved recalls
Follow-up to #3534, from its post-merge review round.
- The abstract-to-overview substitute now applies only to categories whose
stored abstract is the whole file body. A resource or skill whose abstract is
missing (`processing_mode=vectors_only`) or over the per-entry cap read its
body and returned an overview instead, which for a short file is the body
almost verbatim — crossing the opt-in deepening boundary those categories are
documented to have, and doing it even under an explicit `detail="abstract"`.
They now degrade to a bare URI and their body is never read.
- A digest reporting `no_relevant` blanks `rendered`, so the client injects
nothing, yet those URIs still entered the dedup ledger and were cooled for
`dedup_turns` turns. That contradicted the ledger's own bare-URI grace rule
and held memories back from the later turn they were relevant to.
- Flat retrieval reaches built-in memory types outside the four named ones
(`cases`, `patterns`, `tools`, `trajectories`, skill-usage memories) and
reported them as an undeclared `memories` category that no tier or penalty
table covered, so other-peer hits skipped the score penalty and callers could
not pin their tier. The catch-all is now a declared category with both; it
stays out of `quotas`, whose buckets it would overlap. Skill-usage memories
also stop being misread as the `skills` category.
- ZCode, OpenCode and pi own an OV session id but did not forward it, so their
recalls silently ran without query expansion or cross-turn dedup.
- The context-request deadline covered only the server's 30s rewrite fuse, but
the pipeline is serial: expansion, retrieval and budgeting all precede it.
45s covers both fuses and the work between them.
- `plugin` config scope and the `/recall` successor example now match what the
code actually does.
* fix(retrieval): make the context deadline and expansion opt-out reachable
Forwarding a session id turns on server-side query expansion, an LLM call with
its own 5s fuse, but neither the deadline that was supposed to cover it nor the
switch that turns it off reached the two harnesses this PR newly enabled it for.
- `contextRequestTimeoutMs()` now derives the deadline from the request body
rather than from `cfg` plus a rewrite flag. The body is what states which
server stages will run: a session takes the expansion fuse, `rewrite` takes
the digest fuse, and a bare retrieval takes neither and keeps the caller's own
budget. Reading `cfg` alone could not tell those apart.
- OpenCode pinned `timeoutMs: 5000` after spreading the helper's options and pi
ignored them entirely, so the helper's deadline was dead code in both. Their
own budgets are now defaults rather than ceilings. OpenCode's 5s in particular
was shorter than the expansion fuse it had just enabled, so a legal request
would have been aborted client-side and dropped back to the path with neither
dedup nor expansion.
- OpenCode and pi read `OPENVIKING_RECALL_QUERY_EXPANSION` (and
`recallQueryExpansion` in their own config files) and set the `configured`
flag the shared body builder requires, so the documented opt-out exists where
the cost was introduced.
- The integration overview no longer implies every harness reads the same
environment knobs, and describes the deadline as per-stage rather than
rewrite-only.
* feat(retrieval): assemble auto-recall context server-side via /search mode="context"
Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.
This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.
New assembly kernel under openviking/retrieve/context_assembler/:
- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
readable floor, then overview, then full for high-scoring entries. An
oversized tier falls back to the previous one instead of being truncated,
bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
Summary section, code files reuse code_outline signatures, long documents use
a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
directories carry no stored abstract; their full tier stays capped at
overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
every harness inherits cross-turn dedup; exclude_uris remains as the
stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
element per entry. Every tier carries its URI, so the model can always drill
down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
timeout fuses, and a failed rewrite still returns the unrewritten block.
Retrieval failures are counted into stats rather than silently yielding an
empty block.
Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.
* refactor(retrieval): give context tiers a per-category default
The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.
Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.
Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".
Assembly fixes found alongside:
- Removing the abstract cap exemption would turn an oversized abstract into
a bare URI, so it now falls back to overview first — for memory that is a
cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
`asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
such attribute; usage was structurally always null. It now reads the model
instance's tracker and reports only when the call count moved by exactly
one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
fail, and the file was never rewritten, so it could not heal. Records are
now coerced on read and dropped on the next write, along with records left
ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
`viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
resolved a different profile than `POST /recall`; an unknown `detail` value
raised `KeyError` through the whole call instead of degrading.
* feat(codex): inject profile context on session start
Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): raise rewrite timeout default to 30s
* docs(agents): document low-latency recall settings
* fix(codex): prefer luna as recall compressor fallback
* refactor(plugins): unify recall compression setting
* feat(plugins): enable recall compression by default
* docs(agents): use absolute links in image docs
* fix(retrieval): address context assembly review feedback
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test: trim redundant context assembly coverage
* fix(retrieval): address second-round context assembly review
- Drop the backticked `/search` from the deprecated-recall row in both API
overviews. The reference checker scans the whole row after the method cell
for backticked paths, so it read the description as a route named
`POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
belongs to the Rust CLI, which writes `root_api_key`, `output`,
`echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
two Python readers had drifted into stricter subsets, so the shipped example
already failed to load in both. Adding the new `plugin` section to a working
ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
Retrieval validates query and image_url before searching, and the gather
fuse swallowed that rejection along with genuine scope failures, so a body
of `{"mode":"context"}` came back 200 with an empty block instead of the
documented parameter error. Runtime failures still degrade into
`stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
than the 30s fuse, so a rewrite that finished inside its own budget was
aborted client-side, discarding the whole response — including the
uncompressed block the server returns when a rewrite fails — and falling
back to `/recall`. The deadline is only extended when the body actually
requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
`plugin.recallContextTimeoutMs` pins it.
* chore(plugins): sync shared modules into the zcode snapshot
* fix(retrieval): align context quotas and plugin defaults
Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): preserve recall compatibility
Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
Surface the vector store's search_tags on each matched context (returned
under the "tags" key to match the tags filter param) and remove the
result fields the retrieval pipeline never populates (category,
match_reason, relations, overview).
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: add memory plugin mcp harness
* refactor: vendor shared memory plugin modules
* feat: add type quota recall api
* feat: commit codex memory by token threshold
* feat: capture codex tool calls as parts
* feat: add claude skill experience recall
* chore: fix lint in type quota recall server files
* feat: remote marketplace install with unified openviking naming
- Fix root .claude-plugin/marketplace.json git-subdir discriminator key
("type" -> "source"); claude plugin validate now passes.
- Unified installer gains --source remote|archive|dev: remote registers a
synthesized git-subdir marketplace for Claude Code and a git marketplace
for Codex (no repo clone); archive consumes the slim TOS marketplace zip;
dev registers the checkout's examples/ directory for both harnesses.
- One marketplace name (openviking) across all modes and harnesses, so the
plugin id is always openviking-memory@openviking; installer migrates old
openviking-plugins-local registrations and config.toml sections.
- Restore legacy Claude Code (<2.0) support: claude mcp add (stdio proxy)
plus node-based hooks merge into ~/.claude/settings.json.
- Restore optional statusline registration (fetches sources on opt-in).
- Checkbox TUI harness selection via /dev/tty with non-tty fallback.
- Add examples/.agents/plugins/marketplace.json so Codex directory installs
drop the synthetic symlink marketplace.
- Add shared setup wizard (scripts/setup.mjs) for pure-marketplace installs.
- release-tos.yml: upload memory-plugin-shared/install.sh and build/upload
the memory-plugin-marketplace zip; tos-install.sh prefers it and pins all
fetches to TOS via OPENVIKING_SHARED_INSTALL_URL.
- CI: bash -n on installer scripts; marketplace contract tests updated.
* fix(installer): register Claude remote marketplace as a directory
File-type marketplaces (bare marketplace.json path) make Claude Code derive
a wrong installLocation and 'marketplace update' fails with EISDIR. Write
the synthesized manifest to <dir>/.claude-plugin/marketplace.json and add
the directory instead; compare registered sources by exact match so the
old file registration migrates cleanly.
* feat(statusline): show model name and native-style context percentage
A custom statusLine replaces Claude Code's native line including its context
indicator, so reproduce it from the statusline stdin payload: 'Fable 5 ·
ctx 42%' right after the health segment, with native color thresholds
(<70% dim, 70-89% yellow, >=90% red). Falls back from used_percentage to
remaining_percentage to token counts, and stays visible in bypass mode
since it describes the CC conversation, not OV. Opt out with
OPENVIKING_STATUSLINE_CTX=off. Line cap raised 80 -> 100 visible chars.
* fix(installer): keep checkout progress off stdout in plugin_dir_on_disk
Callers capture the function's stdout, so ensure_checkout's info lines were
concatenated into the statusline command registered in settings.json.
* fix(installer): re-register codex git marketplace instead of upgrading
Codex doesn't expose which --ref a git marketplace was added with, and
'marketplace upgrade' refreshes the old ref — so a URL match must not skip
re-registration or a ref override installs the wrong snapshot. Also remove
the stale pre-unification plugin cache directory during migration.
* fix(installer): include .agents in codex sparse checkout
A plugin-dir-only sparse checkout omits the repo-root marketplace manifest
and fails with 'marketplace root does not contain a supported manifest'.
Adding --sparse .agents keeps the snapshot slim (~7.5M vs full repo).
* feat(installer): bilingual prompts, dist channel selection, and TOS git marketplace for codex
- Interactive language selection (English/中文, --lang, auto-detected from
locale); every user-facing prompt is bilingual.
- Download-source selection (--dist github|tos, prompted interactively):
github keeps the remote marketplaces; tos serves GitHub-blocked regions.
- Credentials step now always shows the current ovcli.conf values (masked
key) and offers keep-or-reconfigure instead of silently reusing them.
- Codex on TOS installs from a TOS-hosted git repo over dumb HTTP and keeps
remote updates (codex plugin marketplace upgrade); falls back to the
archive directory if the repo is unavailable. release-tos.yml builds and
uploads the single-commit bare repo (repack + update-server-info).
- Claude Code on TOS warns that directory marketplaces cannot auto-update.
- tos-install.sh bootstraps shrink to TOS_BASE + --dist tos.
- Docs (READMEs, agent-integrations pages, image cards, en+zh) now all use
the single shared installer and drop the deleted wrapper instructions.
* feat(installer): unify all choice prompts on an arrow-key TUI menu
Language, download source, connection mode, keep-or-reconfigure
credentials, statusline enable/replace, and legacy-mode confirmation all
render as the same single-select menu (arrow keys / digit shortcuts /
enter, radio-style highlight) instead of mixed numbered and y/N prompts.
Falls back to numbered input when /dev/tty can't be drawn on and to the
default choice when non-interactive. Free-text fields (URL, API key) stay
line inputs; the harness picker keeps its checkbox multi-select.
* fix(installer): stop piping plugin lists into grep -q under pipefail
grep -q exits on first match and SIGPIPEs the producer, so with pipefail
the 'codex plugin list | grep -q' check read as a miss every time (codex's
list is long; claude's short list masked the bug). Capture the output and
substring-match in bash instead — validation no longer false-warns.
Also: drop the stdio-proxy line from the Done summary; always offer the
install-source menu unless --dist/--source was given (with a checkout the
menu gains a dev option and defaults to it); surface the Claude-on-TOS
no-auto-update warning at source resolution instead of after install.
* fix: unignore examples/memory-plugin-shared/lib and commit the shared modules
The Python build-artifact 'lib/' gitignore rule silently swallowed the
shared plugin module source, so CI checkouts had only the vendored copies
and sync.test.mjs failed with ENOENT on the source directory.
* fix(recall): budget summary/uri fallbacks and sanitize non-finite scores
max_chars is the recall API's contract, but only full fragments counted
toward it — VikingBot's client-side heuristic, faithfully ported, lets
summary and uri fallbacks render far past the budget (repro: max_chars=100
rendered 548 chars). Every fragment now counts; oversized summaries degrade
to uri fragments and entries that can't even fit a uri line are dropped
(reported via stats.dropped). VikingBot itself is intentionally unchanged.
Also run _sanitize_floats over the /recall response like the neighboring
/find and /search routes, so inf/nan scores return 0.0 instead of a 500.
Promote ov_intent_analysis_sft:v7_q8 to the recommended local Ollama
query-planner model. Add the bundled retrieval.ov_intent_analysis_sft_v7
prompt and map v7_q8 to it (v4_q8 mapping kept). Update the setup wizard
presets (v7 recommended, v4 retained, v1 dropped) and the configuration
docs (EN/ZH). Extend tests to cover the v7 mapping.
Co-authored-by: guoxuter <j7azwflq4h@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* Add search retrieval telemetry breakdown
* Fix pyagfs helper annotation imports
* docs: document search relation controls and telemetry fields
Document the new include_relations request parameter and the search telemetry summary fields so the API docs stay aligned with the latest retrieval changes.
* refactor(search): drop relation enrichment and trim telemetry
Remove relation fetching from the retrieval path and delete low-value search telemetry fields so retrieval stays simpler and the telemetry summary focuses on actionable diagnostics.
* feat(cli): add query planner setup to init wizard
Let `openviking-server init` configure the optional lightweight
query_planner model. The wizard pulls the chosen Ollama model and writes
the query_planner config; the IntentAnalyzer selects the matching prompt
at retrieval time via a model->prompt-id mapping, so no prompt files are
copied and no prompts.templates_dir override is needed.
- intent_analyzer: QUERY_PLANNER_PROMPT_BY_MODEL maps the fine-tuned SFT
models to their bundled prompt id; unmapped models keep the default
retrieval.intent_analysis prompt.
- bundle retrieval/ov_intent_analysis_sft_v4.yaml (loaded by its own id).
- ollama detection + doctor now recognize query_planner Ollama usage.
- docs: describe the init flow and runtime prompt selection.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(cli): offer query planner on all paths, recommend it with an Ollama VLM
The init wizard offers the lightweight query planner after model setup. When the
chosen setup already uses an Ollama VLM (`ollama_running` is not None) the
planner rides on that running Ollama at near-zero extra cost, so the enable
prompt is tagged "(recommended)" and defaults to yes. For cloud / non-Ollama VLM
setups it is still offered, but defaults to no and drops the recommendation;
opting in there runs the Ollama install flow.
The Ollama state established during model setup is threaded through the wizard so
the planner reuses it instead of re-running the install dialog:
- `_wizard_ollama` / `_wizard_llamacpp` return `(config, ollama_running)`.
- `run_init` forwards that state to `_wizard_query_planner`.
Docs (zh/en) and tests updated accordingly.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: guoxuter <j7azwflq4h@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Batch child directory vector lookups during recursive retrieval to reduce remote fan-out latency, and accept alternate count aggregate total keys from vector stores.
* feat(session): add account namespace policy and shared sessions
Unify namespace resolution across filesystem, indexing, and session storage.
Add account-shared session paths, role_id auth semantics, and an HTTP demo
script for the four namespace-policy combinations.
* space
* fix(pack): skip derived semantic files in ovpack transfer
Keep ovpack imports resilient to stale sidecars and rebuild semantics through the normal queue instead of restoring derived files verbatim.
* Revert "fix(pack): skip derived semantic files in ovpack transfer"
This reverts commit f4e4db8401.
* fix(namespace): default legacy accounts to agent-shared policy
Clarify that memory.agent_scope_mode is deprecated and document the supported agent memory migration paths.
Introduce native async embedding paths across providers, switch async
retrieval/session hotspots to use them, and add a standalone mixed-load
benchmark plus before/after benchmark evidence for the regression.
* fix: add models observer info for embedder and rerank
* fix: make build deps
* fix: ov observer
* fix: ov observer
---------
Co-authored-by: openviking <openviking@example.com>
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
---------
Co-authored-by: openviking <openviking@example.com>
* feat(embedder): add Cohere dense embedder with embed-v4.0 support
Adds CohereDenseEmbedder using Cohere's Embed API v2.
- Supports embed-v4.0, embed-english-v3.0, embed-multilingual-v3.0
- Server-side dimension reduction for embed-v4.0 (256/512/1024/1536)
- Client-side truncation + renormalization fallback for v3 models
- Asymmetric search via input_type (search_query/search_document)
- Batch embedding with 96-item chunking (Cohere API limit)
- Full factory integration: provider validation, dimension resolution
none
* test(embedder): add unit tests for Cohere embedder
16 tests covering:
- Init validation (api_key required, defaults, model dimensions)
- Dimension handling (v4 server-side, v3 client-side truncation, invalid dims)
- Embedding calls (single, batch, query vs document input_type)
- output_dimension sent for embed-v4.0
- Error handling (API errors → RuntimeError)
- Resource cleanup (close)
none
* feat(rerank): add Cohere rerank-v3.5 support
Extends RerankConfig with provider field and api_key for Cohere.
Adds CohereRerankClient with same interface as VikingDB RerankClient.
HierarchicalRetriever auto-selects rerank backend based on provider.
Config example:
"rerank": {"provider": "cohere", "api_key": "...", "threshold": 0.15}
Quality improvement: META tokenomics query 0.55 → 0.77 relevance score.
none
* test(rerank): add unit tests for Cohere reranker
9 tests covering:
- Rerank batch scoring with index-to-order mapping
- Empty input handling
- API error graceful fallback (returns None)
- Original order preservation from Cohere's sorted response
- Resource cleanup
- RerankConfig provider auto-detection (cohere/vikingdb/empty)
none
* perf(retrieve): increase GLOBAL_SEARCH_TOPK from 5 to 10
More vector candidates for reranker to evaluate = better precision.
With Cohere rerank-v3.5, 10 candidates gives the cross-encoder enough
material to find the best match without excessive latency.
none
* refactor: unify rerank dispatch — route all providers through RerankClient.from_config()
Cohere was special-cased in hierarchical_retriever.py while openai/litellm
went through the centralized RerankClient.from_config() dispatch. This commit
adds CohereRerankClient.from_config() and routes it through the same path.
Also fixes a bug where from_config() used config.provider directly instead
of _effective_provider(), which meant auto-detected providers (e.g. api_key
without explicit provider="cohere") would not dispatch correctly.
none
The C++ vector engine can produce infinity values from inner product
overflow on non-normalized vectors (distance_type=ip with
NormalizeVector=False). These inf scores propagate through the adapter
and retriever layers, ultimately crashing JSON serialization with
ValueError.
Fix applied at two layers:
1. CollectionAdapter.query() — clamp non-finite scores to 0.0
immediately after reading from the C++ engine result, before
they enter the JSON-serializable record chain.
2. HierarchicalRetriever — clamp _score values at all three entry
points (_merge_starting_points, _prepare_initial_candidates,
_recursive_search) before score propagation arithmetic can
amplify inf into downstream computations.
The existing isfinite guard in _convert_to_matched_contexts catches
scores only at the final conversion step, which is too late — inf
values already cause serialization failures in intermediate API
responses and vectordb service endpoints.
Closes#871
When local vector search returns inf scores (e.g., zero vectors or
embedding overflow), the hierarchical retriever passes them through to
the API response. FastAPI/Starlette's JSON encoder rejects inf/nan with:
ValueError: Out of range float values are not JSON compliant: inf
Fix:
1. hierarchical_retriever.py: clamp semantic_score and final_score to 0.0
when math.isfinite() returns False
2. search.py: add _sanitize_floats() as a defense-in-depth layer on the
find and search endpoints
Closes #inf-score
Co-authored-by: a1461750564 <a1461750564@users.noreply.github.com>
* feat(telemetry): add Prometheus metrics exporter via observer pattern
Adds PrometheusObserver implementing BaseObserver with thread-safe
counters and histograms for retrieval, embedding, VLM, and cache
metrics. Exposes /metrics endpoint in Prometheus text exposition
format. Opt-in via server.telemetry.prometheus.enabled config.
No new dependencies - generates Prometheus text format manually.
* style: use dict.fromkeys per ruff C420
* fix(telemetry): wire PrometheusObserver into data collection and address review feedback
- Hook observer into RetrievalStatsCollector and other data paths
- Register metrics router statically in create_app()
- Remove unrelated with_bot/bot_api_url config changes
* fix(telemetry): measure VLM call duration at call sites
Time each VLM API call using time.perf_counter() and pass the
measured duration through to update_token_usage(), which records
it in the Prometheus histogram.
Previously duration_seconds always defaulted to 0.0 because no
backend passed actual timing data. Now all three backends (OpenAI,
VolcEngine, LiteLLM) measure wall-clock time around the API call
in get_completion, get_completion_async, get_vision_completion,
and get_vision_completion_async.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Add retrieval quality observability following the existing observer
pattern (QueueObserver, VLMObserver, VikingDBObserver).
- RetrievalStatsCollector: thread-safe singleton that accumulates
per-query metrics (result counts, scores, latency, rerank usage)
- RetrievalObserver: reads accumulated stats, reports health based on
zero-result rate, formats status table with tabulate
- Instrumented HierarchicalRetriever.retrieve() to record stats
- Added /api/v1/observer/retrieval endpoint
- Included in system-wide observer health check
- 17 unit tests covering stats, collector, and observer
This contribution was developed with AI assistance (Claude Code).
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
* fix: windows zip path norm
* fix: account id in vector db
* fix: add some log
* fix: add some log, and fixed search
* fix: add some log, and fixed search
* fix: add some log, and fixed search
* fix: add some log, and fixed search
* fix: add some log, and fixed search
* fix: add some log, and fixed search
* fix: add some log, and fixed search
* fix: add some log, and fixed search
---------
Co-authored-by: openviking <openviking@example.com>