Commit Graph
58 Commits
Author SHA1 Message Date
Qin Haojie 9295a3b955 fix(session): remove actor scope from session lifecycle (#3661)
Keep sessions user-scoped and remove legacy agent fallback that could leak an actor view into commit memory writes.
2026-07-31 19:52:19 +08:00
baojun-zhang a1b6626558 fix(import): avoid runtime storage package lazy imports for VikingDB managers (#3616) 2026-07-29 21:29:13 +08:00
19ca274a24 fix(retrieve): bound reranker input size (#3289)
* fix(retrieve): bound reranker input size

* fix(retrieve): make rerank input limit opt-in

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-16 17:19:38 +08:00
huangruitengandhuangruiteng d14c0673ca fix(recall): hide memory fields metadata (#3240)
Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-14 17:41:30 +08:00
huangruitengandhuangruiteng ba46491af0 fix(retrieve): preserve rerank fallbacks for empty documents (#3231)
Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-14 11:51:35 +08:00
huangruitengandhuangruiteng 90b2c910d7 perf: parallelize type quota recall searches (#3175)
Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-13 13:50:50 +08:00
t0saki 2a81edc707 Add workspace peer mode for memory plugins (#3099) 2026-07-09 17:53:36 +08:00
Qin Haojie 3003ed61d7 feat(retrieval): support image search (#3093)
Add multimodal image vectorization and image query support across the server, SDKs, and CLI.
2026-07-09 16:42:53 +08:00
Qin Haojie 7bbf6371f8 fix(retrieval): offload rerank calls from event loop (#3063)
Run sync rerank providers in the thread pool so slow rerank requests do not block health checks on single-worker servers.
2026-07-07 17:17:18 +08:00
t0saki f905562534 feat(plugins): stdio MCP proxy, remote marketplace install, and type-quota recall for memory plugins (#3039)
* feat: add memory plugin mcp harness

* refactor: vendor shared memory plugin modules

* feat: add type quota recall api

* feat: commit codex memory by token threshold

* feat: capture codex tool calls as parts

* feat: add claude skill experience recall

* chore: fix lint in type quota recall server files

* feat: remote marketplace install with unified openviking naming

- Fix root .claude-plugin/marketplace.json git-subdir discriminator key
  ("type" -> "source"); claude plugin validate now passes.
- Unified installer gains --source remote|archive|dev: remote registers a
  synthesized git-subdir marketplace for Claude Code and a git marketplace
  for Codex (no repo clone); archive consumes the slim TOS marketplace zip;
  dev registers the checkout's examples/ directory for both harnesses.
- One marketplace name (openviking) across all modes and harnesses, so the
  plugin id is always openviking-memory@openviking; installer migrates old
  openviking-plugins-local registrations and config.toml sections.
- Restore legacy Claude Code (<2.0) support: claude mcp add (stdio proxy)
  plus node-based hooks merge into ~/.claude/settings.json.
- Restore optional statusline registration (fetches sources on opt-in).
- Checkbox TUI harness selection via /dev/tty with non-tty fallback.
- Add examples/.agents/plugins/marketplace.json so Codex directory installs
  drop the synthetic symlink marketplace.
- Add shared setup wizard (scripts/setup.mjs) for pure-marketplace installs.
- release-tos.yml: upload memory-plugin-shared/install.sh and build/upload
  the memory-plugin-marketplace zip; tos-install.sh prefers it and pins all
  fetches to TOS via OPENVIKING_SHARED_INSTALL_URL.
- CI: bash -n on installer scripts; marketplace contract tests updated.

* fix(installer): register Claude remote marketplace as a directory

File-type marketplaces (bare marketplace.json path) make Claude Code derive
a wrong installLocation and 'marketplace update' fails with EISDIR. Write
the synthesized manifest to <dir>/.claude-plugin/marketplace.json and add
the directory instead; compare registered sources by exact match so the
old file registration migrates cleanly.

* feat(statusline): show model name and native-style context percentage

A custom statusLine replaces Claude Code's native line including its context
indicator, so reproduce it from the statusline stdin payload: 'Fable 5 ·
ctx 42%' right after the health segment, with native color thresholds
(<70% dim, 70-89% yellow, >=90% red). Falls back from used_percentage to
remaining_percentage to token counts, and stays visible in bypass mode
since it describes the CC conversation, not OV. Opt out with
OPENVIKING_STATUSLINE_CTX=off. Line cap raised 80 -> 100 visible chars.

* fix(installer): keep checkout progress off stdout in plugin_dir_on_disk

Callers capture the function's stdout, so ensure_checkout's info lines were
concatenated into the statusline command registered in settings.json.

* fix(installer): re-register codex git marketplace instead of upgrading

Codex doesn't expose which --ref a git marketplace was added with, and
'marketplace upgrade' refreshes the old ref — so a URL match must not skip
re-registration or a ref override installs the wrong snapshot. Also remove
the stale pre-unification plugin cache directory during migration.

* fix(installer): include .agents in codex sparse checkout

A plugin-dir-only sparse checkout omits the repo-root marketplace manifest
and fails with 'marketplace root does not contain a supported manifest'.
Adding --sparse .agents keeps the snapshot slim (~7.5M vs full repo).

* feat(installer): bilingual prompts, dist channel selection, and TOS git marketplace for codex

- Interactive language selection (English/中文, --lang, auto-detected from
  locale); every user-facing prompt is bilingual.
- Download-source selection (--dist github|tos, prompted interactively):
  github keeps the remote marketplaces; tos serves GitHub-blocked regions.
- Credentials step now always shows the current ovcli.conf values (masked
  key) and offers keep-or-reconfigure instead of silently reusing them.
- Codex on TOS installs from a TOS-hosted git repo over dumb HTTP and keeps
  remote updates (codex plugin marketplace upgrade); falls back to the
  archive directory if the repo is unavailable. release-tos.yml builds and
  uploads the single-commit bare repo (repack + update-server-info).
- Claude Code on TOS warns that directory marketplaces cannot auto-update.
- tos-install.sh bootstraps shrink to TOS_BASE + --dist tos.
- Docs (READMEs, agent-integrations pages, image cards, en+zh) now all use
  the single shared installer and drop the deleted wrapper instructions.

* feat(installer): unify all choice prompts on an arrow-key TUI menu

Language, download source, connection mode, keep-or-reconfigure
credentials, statusline enable/replace, and legacy-mode confirmation all
render as the same single-select menu (arrow keys / digit shortcuts /
enter, radio-style highlight) instead of mixed numbered and y/N prompts.
Falls back to numbered input when /dev/tty can't be drawn on and to the
default choice when non-interactive. Free-text fields (URL, API key) stay
line inputs; the harness picker keeps its checkbox multi-select.

* fix(installer): stop piping plugin lists into grep -q under pipefail

grep -q exits on first match and SIGPIPEs the producer, so with pipefail
the 'codex plugin list | grep -q' check read as a miss every time (codex's
list is long; claude's short list masked the bug). Capture the output and
substring-match in bash instead — validation no longer false-warns.

Also: drop the stdio-proxy line from the Done summary; always offer the
install-source menu unless --dist/--source was given (with a checkout the
menu gains a dev option and defaults to it); surface the Claude-on-TOS
no-auto-update warning at source resolution instead of after install.

* fix: unignore examples/memory-plugin-shared/lib and commit the shared modules

The Python build-artifact 'lib/' gitignore rule silently swallowed the
shared plugin module source, so CI checkouts had only the vendored copies
and sync.test.mjs failed with ENOENT on the source directory.

* fix(recall): budget summary/uri fallbacks and sanitize non-finite scores

max_chars is the recall API's contract, but only full fragments counted
toward it — VikingBot's client-side heuristic, faithfully ported, lets
summary and uri fallbacks render far past the budget (repro: max_chars=100
rendered 548 chars). Every fragment now counts; oversized summaries degrade
to uri fragments and entries that can't even fit a uri line are dropped
(reported via stats.dropped). VikingBot itself is intentionally unchanged.

Also run _sanitize_floats over the /recall response like the neighboring
/find and /search routes, so inf/nan scores return 0.0 instead of a 500.
2026-07-07 12:33:59 +08:00
Qin Haojie 93dbd223ab refactor(retrieval): simplify quick search flow (#2812) 2026-06-26 14:50:05 +08:00
9506101bd3 feat(retrieval): recommend ov_intent_analysis_sft v7_q8 query planner (#2624)
Promote ov_intent_analysis_sft:v7_q8 to the recommended local Ollama
query-planner model. Add the bundled retrieval.ov_intent_analysis_sft_v7
prompt and map v7_q8 to it (v4_q8 mapping kept). Update the setup wizard
presets (v7 recommended, v4 retained, v1 dropped) and the configuration
docs (EN/ZH). Extend tests to cover the v7 mapping.

Co-authored-by: guoxuter <j7azwflq4h@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 21:30:35 +08:00
Jiahui Zhou 27e90ad9cd Feature/search telemetry relations control (#2547)
* Add search retrieval telemetry breakdown

* Fix pyagfs helper annotation imports

* docs: document search relation controls and telemetry fields

Document the new include_relations request parameter and the search telemetry summary fields so the API docs stay aligned with the latest retrieval changes.

* refactor(search): drop relation enrichment and trim telemetry

Remove relation fetching from the retrieval path and delete low-value search telemetry fields so retrieval stays simpler and the telemetry summary focuses on actionable diagnostics.
2026-06-10 20:00:47 +08:00
1c0bbcdf0e feat(cli): add query planner setup to init wizard (#2485)
* feat(cli): add query planner setup to init wizard

Let `openviking-server init` configure the optional lightweight
query_planner model. The wizard pulls the chosen Ollama model and writes
the query_planner config; the IntentAnalyzer selects the matching prompt
at retrieval time via a model->prompt-id mapping, so no prompt files are
copied and no prompts.templates_dir override is needed.

- intent_analyzer: QUERY_PLANNER_PROMPT_BY_MODEL maps the fine-tuned SFT
  models to their bundled prompt id; unmapped models keep the default
  retrieval.intent_analysis prompt.
- bundle retrieval/ov_intent_analysis_sft_v4.yaml (loaded by its own id).
- ollama detection + doctor now recognize query_planner Ollama usage.
- docs: describe the init flow and runtime prompt selection.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(cli): offer query planner on all paths, recommend it with an Ollama VLM

The init wizard offers the lightweight query planner after model setup. When the
chosen setup already uses an Ollama VLM (`ollama_running` is not None) the
planner rides on that running Ollama at near-zero extra cost, so the enable
prompt is tagged "(recommended)" and defaults to yes. For cloud / non-Ollama VLM
setups it is still offered, but defaults to no and drops the recommendation;
opting in there runs the Ollama install flow.

The Ollama state established during model setup is threaded through the wizard so
the planner reuses it instead of re-running the install dialog:

- `_wizard_ollama` / `_wizard_llamacpp` return `(config, ollama_running)`.
- `run_init` forwards that state to `_wizard_query_planner`.

Docs (zh/en) and tests updated accordingly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: guoxuter <j7azwflq4h@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 17:12:05 +08:00
Qin Haojie ff258768c2 feat(memory): 引入 User/Peer 记忆隔离模型 (#2236)
* feat(memory): introduce user and peer memory isolation

Unify agent-scoped memory behavior into user-owned memory spaces, add peer_id compatibility for session and retrieval paths, and wire memory_policy through session commit flows.

* feat(memory): align session identity around peer IDs

* feat(search): pass peer id through retrieval

* refactor(memory): remove agent identity from integrations

* fix(memory): isolate peer identity from self extraction

* fix(tau2): provision benchmark user configs

* fix(auth): allow admin keys to access data APIs

* fix(openclaw): enable peer memory policy for peer roles

* fix(openclaw): resolve sender for peer recall

* refactor(session): simplify memory extraction routing

* refactor(ov-cli): reduce formatting-only diff

* refactor(message): remove unused message helpers

* refactor(retrieval): simplify peer target resolution

* refactor(namespace): remove deprecated agent namespace policy

* fix(agent): propagate peer id through integrations

* fix(auth): align integration clients with api-key mode
2026-06-05 10:55:48 +08:00
agent 2b5fb5bd7a feat(search) Add lightweight query planner config for intent analysis (#2224)
* feat: query planer

* feat: add planer

* docs: clarify optional query planner config

* refactor: centralize query planner selection

* feat: file

* fix: format file
2026-05-27 16:30:28 +08:00
Qin Haojie ba9df59f95 perf(retrieval): parallelize hierarchical child search (#2155)
Batch child directory vector lookups during recursive retrieval to reduce remote fan-out latency, and accept alternate count aggregate total keys from vector stores.
2026-05-21 13:56:43 +08:00
zgy 79389ee474 feat(retriever): add level filter to find with retriever-level filtering, preserving recursive navigation (#1988)
* Feat(retriever): add level filter to find with retriever-level filtering, preserving recursive navigation

Add level: Optional[List[int]] parameter to find API/CLI/SDK to filter
results by L0 (abstract), L1 (overview), L2 (original file).

Key design: filter at two result collection points inside
HierarchicalRetriever (global search pool + recursive traversal pool),
NOT at vector search layer or post-filter layer. This preserves L0/L1
directory waypoints in dir_queue for recursive navigation, avoiding the
quality regression that merge_level_filter (PR #1980) causes.

Changes:
- Python: FindRequest, SearchService, VikingFS, LocalClient,
  AsyncOpenViking, SyncOpenViking, MCP endpoint all pass level through
- Retriever: initial_candidates and collected_by_uri filtered by level;
  dir_queue navigation unchanged
- Fix variable shadowing: rename debug loop 'level' to 'result_level'
- Add stagnation detection to convergence check when level filter
  prevents reaching limit
- CLI: --level / -L flag with Option<Vec<i32>> + value_delimiter
- Remove merge_level_filter from find route (breaks recursive navigation)
- Tests: 8 new test cases covering param passthrough, backward compat,
  single/mixed level filtering, and edge cases

* feat(search): 为 search 方法添加 level 过滤,与 find 保持一致

- HTTP 路由层:FindRequest.level 改为 Union[int, str, List[int]],search 路由传递 level 参数
- Service 层:SearchService.search 增加 level 参数
- 存储层:VikingFS.search 增加 level 参数,传递给 retriever.retrieve
- SDK 层:LocalClient/AsyncOpenViking/SyncOpenViking.search 增加 level 参数
- MCP 端点:search 工具增加 level 参数
- CLI 层:ov search --level 从 Option<String> 改为 Option<Vec<i32>>,与 find 一致
- 删除 append_level_filter_params,改用内联格式化
- 路由层用 _resolve_levels() 统一将 Union 类型转为 List[int]
- 新增 9 个测试:7 个 search level 测试 + 2 个 find Union 类型输入测试
2026-05-14 18:06:03 +08:00
Monday 5154644c01 fix(retrieve): apply threshold filter to initial candidates in HierarchicalRetriever (#1797) 2026-04-29 17:01:11 +08:00
Qin Haojie b35d38a323 feat(config): 配置检索打分和 embedding 输入 (#1770)
* feat(retrieval): configure hotness score blending

* feat(retrieval): configure score propagation alpha

* test(retrieval): trim redundant propagation coverage

* feat(embedding): centralize token estimation

* fix(embedding): use shared token estimator

* fix(embedding): narrow token truncation scope
2026-04-28 19:04:24 +08:00
Qin Haojie cebc45907b feat(session): add account namespace policy and shared sessions (#1356)
* feat(session): add account namespace policy and shared sessions

Unify namespace resolution across filesystem, indexing, and session storage.
Add account-shared session paths, role_id auth semantics, and an HTTP demo
script for the four namespace-policy combinations.

* space

* fix(pack): skip derived semantic files in ovpack transfer

Keep ovpack imports resilient to stale sidecars and rebuild semantics through the normal queue instead of restoring derived files verbatim.

* Revert "fix(pack): skip derived semantic files in ovpack transfer"

This reverts commit f4e4db8401.

* fix(namespace): default legacy accounts to agent-shared policy

Clarify that memory.agent_scope_mode is deprecated and document the supported agent memory migration paths.
2026-04-17 15:12:45 +08:00
baojun-zhang b441622ee6 feat(metric): add metric system (#1357)
* feat(metric): add metric system

* feat(metric): add metric system

* feat(metric): add metric system

* fix(metric): do not cancel refresh tasks on deadline; trust only authenticated account id; avoid per-scrape rerank clients

* doc(metric): add metric guide doc

* doc(metric): add metric guide doc

* doc(metric): add metric guide doc

* doc(metric): add metric guide doc

* feat(metric): fix bug & change account dimension switch to default true
2026-04-14 20:55:02 +08:00
Qin Haojie 804de2a2ad fix(embedder): reduce async contention in session flows (#1301)
Introduce native async embedding paths across providers, switch async
retrieval/session hotspots to use them, and add a standalone mixed-load
benchmark plus before/after benchmark evidence for the regression.
2026-04-08 16:31:14 +08:00
Jiahui Zhou fea7c01ed5 Revert "feat(retrieve): use tags metadata for cross-subtree retrieval (#1162)" (#1200)
This reverts commit e72b614b3d.
2026-04-03 14:37:47 +08:00
13ernkastel e72b614b3d feat(retrieve): use tags metadata for cross-subtree retrieval (#1162)
* feat: use tags to expand cross-subtree retrieval

* fix: harden sync retrieval argument forwarding

* fix: resolve PR lint failures

* feat(tags): namespace stored and queried resource tags

* fix(tags): enforce canonical tag namespaces

* style: sort resource service imports
2026-04-03 13:24:38 +08:00
MaojiaShengandopenviking d739f742d7 fix: ov status shows embedding and rerank models usage (#1191)
* fix: add models observer info for embedder and rerank

* fix: make build deps

* fix: ov observer

* fix: ov observer

---------

Co-authored-by: openviking <openviking@example.com>
2026-04-03 09:00:38 +08:00
MaojiaShengandopenviking ce998873f9 lisence: change the main lisence to AGPL-3.0 (#1085)
* lisence: change the main lisence from Apache-2.0 to AGPL-v3

* lisence: change the main lisence from Apache-2.0 to AGPL-v3

* lisence: change the main lisence from Apache-2.0 to AGPL-v3

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-30 14:37:42 +08:00
Dico Angelo 9bcdaf473e feat(embedder): add Cohere dense embedder (#941)
* feat(embedder): add Cohere dense embedder with embed-v4.0 support

Adds CohereDenseEmbedder using Cohere's Embed API v2.
- Supports embed-v4.0, embed-english-v3.0, embed-multilingual-v3.0
- Server-side dimension reduction for embed-v4.0 (256/512/1024/1536)
- Client-side truncation + renormalization fallback for v3 models
- Asymmetric search via input_type (search_query/search_document)
- Batch embedding with 96-item chunking (Cohere API limit)
- Full factory integration: provider validation, dimension resolution

none

* test(embedder): add unit tests for Cohere embedder

16 tests covering:
- Init validation (api_key required, defaults, model dimensions)
- Dimension handling (v4 server-side, v3 client-side truncation, invalid dims)
- Embedding calls (single, batch, query vs document input_type)
- output_dimension sent for embed-v4.0
- Error handling (API errors → RuntimeError)
- Resource cleanup (close)

none

* feat(rerank): add Cohere rerank-v3.5 support

Extends RerankConfig with provider field and api_key for Cohere.
Adds CohereRerankClient with same interface as VikingDB RerankClient.
HierarchicalRetriever auto-selects rerank backend based on provider.

Config example:
  "rerank": {"provider": "cohere", "api_key": "...", "threshold": 0.15}

Quality improvement: META tokenomics query 0.55 → 0.77 relevance score.

none

* test(rerank): add unit tests for Cohere reranker

9 tests covering:
- Rerank batch scoring with index-to-order mapping
- Empty input handling
- API error graceful fallback (returns None)
- Original order preservation from Cohere's sorted response
- Resource cleanup
- RerankConfig provider auto-detection (cohere/vikingdb/empty)

none

* perf(retrieve): increase GLOBAL_SEARCH_TOPK from 5 to 10

More vector candidates for reranker to evaluate = better precision.
With Cohere rerank-v3.5, 10 candidates gives the cross-encoder enough
material to find the best match without excessive latency.

none

* refactor: unify rerank dispatch — route all providers through RerankClient.from_config()

Cohere was special-cased in hierarchical_retriever.py while openai/litellm
went through the centralized RerankClient.from_config() dispatch. This commit
adds CohereRerankClient.from_config() and routes it through the same path.

Also fixes a bug where from_config() used config.provider directly instead
of _effective_provider(), which meant auto-detected providers (e.g. api_key
without explicit provider="cohere") would not dispatch correctly.

none
2026-03-30 14:15:08 +08:00
Sergey Nemesh 2a93185397 fix: clamp inf/nan similarity scores at source to prevent JSON crash (#882)
The C++ vector engine can produce infinity values from inner product
overflow on non-normalized vectors (distance_type=ip with
NormalizeVector=False). These inf scores propagate through the adapter
and retriever layers, ultimately crashing JSON serialization with
ValueError.

Fix applied at two layers:

1. CollectionAdapter.query() — clamp non-finite scores to 0.0
   immediately after reading from the C++ engine result, before
   they enter the JSON-serializable record chain.

2. HierarchicalRetriever — clamp _score values at all three entry
   points (_merge_starting_points, _prepare_initial_candidates,
   _recursive_search) before score propagation arithmetic can
   amplify inf into downstream computations.

The existing isfinite guard in _convert_to_matched_contexts catches
scores only at the final conversion step, which is too late — inf
values already cause serialization failures in intermediate API
responses and vectordb service endpoints.

Closes #871
2026-03-23 11:31:21 +08:00
月球宇航员anda1461750564 d1f84235e5 fix(search): clamp inf/nan scores from vector search to prevent JSON serialization failure (#824)
When local vector search returns inf scores (e.g., zero vectors or
embedding overflow), the hierarchical retriever passes them through to
the API response. FastAPI/Starlette's JSON encoder rejects inf/nan with:
  ValueError: Out of range float values are not JSON compliant: inf

Fix:
1. hierarchical_retriever.py: clamp semantic_score and final_score to 0.0
   when math.isfinite() returns False
2. search.py: add _sanitize_floats() as a defense-in-depth layer on the
   find and search endpoints

Closes #inf-score

Co-authored-by: a1461750564 <a1461750564@users.noreply.github.com>
2026-03-21 14:36:45 +08:00
08d9949072 feat(telemetry): add Prometheus metrics exporter via observer pattern (#806)
* feat(telemetry): add Prometheus metrics exporter via observer pattern

Adds PrometheusObserver implementing BaseObserver with thread-safe
counters and histograms for retrieval, embedding, VLM, and cache
metrics. Exposes /metrics endpoint in Prometheus text exposition
format. Opt-in via server.telemetry.prometheus.enabled config.

No new dependencies - generates Prometheus text format manually.

* style: use dict.fromkeys per ruff C420

* fix(telemetry): wire PrometheusObserver into data collection and address review feedback

- Hook observer into RetrievalStatsCollector and other data paths
- Register metrics router statically in create_app()
- Remove unrelated with_bot/bot_api_url config changes

* fix(telemetry): measure VLM call duration at call sites

Time each VLM API call using time.perf_counter() and pass the
measured duration through to update_token_usage(), which records
it in the Prometheus histogram.

Previously duration_seconds always defaulted to 0.0 because no
backend passed actual timing data. Now all three backends (OpenAI,
VolcEngine, LiteLLM) measure wall-clock time around the API call
in get_completion, get_completion_async, get_vision_completion,
and get_vision_completion_async.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-03-21 14:20:38 +08:00
Jiahui Zhou 0bd5b1df73 fix(recall): fix recall limit (#821) 2026-03-20 19:31:24 +08:00
mildred522 82307c3281 fix(retrieval): allow find without rerank and preserve level-2 rerank scores (#754)
* fix: allow find without rerank config

* fix: preserve rerank scores for initial candidates
2026-03-19 10:42:05 +08:00
Xiaogang Zhouandxiaogang.zhou 9800b4001b feat(embedding): simplify the asymmetric embedder (#702)
Co-authored-by: xiaogang.zhou <xiaogang.zhou@bytedance.com>
2026-03-18 00:18:44 +08:00
zhoujiahui b280b56b30 feat(trace): add request-level trace metrics and API support (#640)
refactor: replace operation trace with telemetry

fix telemetry demo skill ingestion

simplify telemetry summary metric keys

rename remaining trace telemetry artifacts

feat: support configurable telemetry payloads

docs: rewrite operation telemetry design in chinese

fix: reject telemetry for async session commit

refactor: isolate telemetry orchestration

refactor: remove telemetry from find payloads

refactor: remove telemetry event payloads

fix(trace): keep only telemetry-related changes

fix(trace): remove top-level usage from telemetry responses

feat(console): default telemetry on proxied operations
2026-03-15 22:44:15 +08:00
Matt Van HornandMatt Van Horn c44e6147c1 feat(retrieve): add RetrievalObserver for retrieval quality metrics (#622)
Add retrieval quality observability following the existing observer
pattern (QueueObserver, VLMObserver, VikingDBObserver).

- RetrievalStatsCollector: thread-safe singleton that accumulates
  per-query metrics (result counts, scores, latency, rerank usage)
- RetrievalObserver: reads accumulated stats, reports health based on
  zero-result rate, formats status table with tabulate
- Instrumented HierarchicalRetriever.retrieve() to record stats
- Added /api/v1/observer/retrieval endpoint
- Included in system-wide observer health check
- 17 unit tests covering stats, collector, and observer

This contribution was developed with AI assistance (Claude Code).

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
2026-03-15 14:51:21 +08:00
MaojiaShengandopenviking 0fb8f13a13 fix: windows zip path, code repo indexing, search retrieval, account id, rust cli version... (#577)
* fix: windows zip path norm

* fix: account id in vector db

* fix: add some log

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-15 10:04:51 +08:00
mildred522 73c7072429 fix: integrate rerank into hierarchical retriever (#599)
* feat(retrieval): integrate rerank into hierarchical retriever

* chore(retrieval): align retriever typing with lint checks

* docs(retrieval): clarify rerank fallback in zh docs
2026-03-14 17:46:41 +08:00
lixingjia c80a2e45e5 fix(search): use session summaries in search and cap intent summary length (#494) 2026-03-09 20:37:47 +08:00
zztdan 62d479bd53 fix: improve ISO datetime parsing (#404) 2026-03-04 11:30:50 +08:00
zhoujiahui 2699d7c0de Fix retriever (#382)
* fix: fix retriever

* fix: fix recall
2026-03-02 21:47:42 +08:00
kkkwjx e8981bf879 Feat/vectordb interface refactor (#327)
* refactor: route vector access through semantic gateway

* refactor vector storage to collection-bound drivers

* refactor(storage): collapse gateway/interface into single-collection backend

* refactor vectordb backend to single-collection adapter model

* chore: align naming with vikingdb and rename session test

* fix

* docs: add guide for integrating third-party vectordb adapters
2026-02-27 15:56:07 +08:00
chuanbao666 4c9340fc10 feat(agfs): agfs新增binding client (#304)
* feat(agfs): agfs新增binding client

fix: viking_fs适配binding client, 简化agfs相关参数

* feat(agfs): agfs binding-client支持windows

* docs: 补充agfs binding-client用法

* fix(agfs): 默认带上agfs binding-client依赖库; setup.py支持编译该库

* fix(agfs): add agfs binding-client linux so

* fix(agfs): code format
2026-02-26 20:27:00 +08:00
r266-techandr266-tech a88207f9c3 feat(memory): add hotness scoring for cold/hot memory lifecycle (#296) (#297)
Add a hotness_score() function that combines access frequency (sigmoid of
log1p(active_count)) with time-based recency decay (exponential, 7-day
half-life) to produce a 0.0-1.0 score for each context.

The hierarchical retriever now blends this hotness score with the semantic
similarity score: final = (1 - alpha) * semantic + alpha * hotness, where
alpha defaults to 0.2 (HOTNESS_ALPHA class constant).

Key properties:
- alpha=0 preserves existing behavior exactly (backward compatible)
- Default alpha=0.2 gives a mild boost without overriding semantic relevance
- No schema changes, no Context class modifications
- Pure additive: new file + minimal retriever modification

Closes #296

Co-authored-by: r266-tech <r266-tech@users.noreply.github.com>
2026-02-26 16:55:25 +08:00
kkkwjx 5982d95977 refactor(filter): replace prefix operator with must (#292) 2026-02-25 21:04:53 +08:00
kkkwjxandqin-ctx 37b5d0d359 Multi tenant (#283)
* test: add multi-tenant isolation unit coverage

* feat: implement phase2 multi-tenant context isolation

* feat: multi tenant

* fix: resolve multi-tenant merge conflicts

* feat: multi tenant

* fix: enforce tenant isolation for session get

* tos config

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-02-25 18:25:36 +08:00
MaojiaShengandopenviking a412c7d164 feat: break change, remove is_leaf scalar and use level instead (#271)
* feat: wrap log configs in LogConfig, add log rotation, fix empty file creation
- Create new LogConfig class to wrap all log-related configuration
- Update OpenVikingConfig to use nested log config instead of separate fields
- Add log rotation support with TimedRotatingFileHandler in logger.py
- Add configure_uvicorn_logging to make uvicorn use OpenViking's logging config
- Fix empty file being created in current directory by properly handling log.output=file in logging_init.py
- Update examples/ov.conf.example to use new nested log config structure
- Update __init__.py to export LogConfig and initialize_openviking_config

* docs: help to configure log and workspace

* feat: break change, remove is_leaf scalar and use level instead

* feat: break change, remove is_leaf scalar and use level instead

* feat: rust cli add-resource support zip

---------

Co-authored-by: openviking <openviking@example.com>
2026-02-25 13:01:41 +08:00
chuanbao666 ec7afc6e51 增加openviking/eval模块,用于评估测试 (#265)
* feat(eval): 增加评估模块,对viking_fs s3后端测试

* fix: 修复s3测试

* fix(eval): fix ruff check

* refactor: eval细分ragas模块用于rag相关评测
2026-02-24 17:37:10 +08:00
Zayn Jarvis 3f65a326eb fix: skill search ranking - use overview for embedding and fix visited set filtering (#228)
1. skill_processor.py: Use LLM-generated overview for vectorization
   instead of short abstract, aligning with how resources handle
   directory vectorization in semantic_processor.py.

2. hierarchical_retriever.py: Separate 'visited for traversal' from
   'collected as result'. The visited set previously dropped the most
   relevant results - global search found them first, marked visited,
   then parent directory search skipped them as children.
2026-02-20 12:17:02 +08:00
mildred522 ecc1ada2a7 fix: target directories retrieve (#227)
* perf: reuse query embeddings in hierarchical retriever

* fix(retrieve): honor target_directories filtering
2026-02-20 12:15:30 +08:00