Commit Graph
26 Commits
Author SHA1 Message Date
Jiahui ZhouandTRAE CLI f857349d6b feat(vectordb): add text field type for large strings (#3725)
The bytes_row STRING type uses a uint16 length prefix, capping a single
field at 65535 bytes. Add a new TEXT field type (enum value 9) that mirrors
STRING semantics (utf-8 str round-trip) but uses a uint32 length prefix,
lifting the per-field limit to ~4GB.

TEXT is added only at the physical bytes_row layer, across all serializers
that must stay byte-identical: the C++ engine (bytes_row.h/.cpp), the abi3
boundary (abi3_engine_backend.cpp, decoding to str not bytes), the pure
Python fallback (store/bytes_row.py), and the engine API (_python_api.py).
Existing types and the CandidateData.fields field are untouched, so old
on-disk data stays readable without reindex.

Fields opt into the new type via metadata={"field_type": FieldType.text}.

Add TestTextFieldType covering >65535-byte round-trips, py<->cpp cross
read/write, binary consistency, and metadata-based declaration.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-04 11:33:18 +08:00
Yuanqing ZHAOandYuanqing Zhao 40dd05271c perf(vectordb): micro-batch compatible cuVS searches (#3382)
* perf(vectordb): micro-batch compatible cuVS searches

* fix(vectordb): serialize micro-batch device admission

* perf(vectordb): pipeline warm cuVS micro-batch admission

* docs(cuvs): align micro-batching guidance

* fix(cuvs): warm-batch empty filters

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-23 11:10:52 +08:00
Yuanqing ZHAOandYuanqing Zhao 2d84a6f13c perf(cuvs): improve recovery scalability and filtered-search efficiency (#3311)
* perf(index): avoid materializing descendant path strings

* perf(cuvs): compact host vector shadow

* perf(storage): prune redundant tenant path scopes

* fix(cuvs): synchronize worker-thread searches

* perf(cuvs): stream dense shadow recovery

* test(storage): cover cross-user path scopes

* perf(storage): page candidate recovery scans

* perf(cuvs): accelerate adaptive filter routing

* perf(cuvs): reduce filtered-search host overhead

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-21 15:59:58 +08:00
Yuanqing ZHAOandYuanqing Zhao e1cea0998b perf(vectordb): avoid hydrating unrequested vectors (#3375)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-20 19:13:00 +08:00
Yuanqing ZHAOandYuanqing Zhao fa19ac0a75 perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest (#3277)
* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest

Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.

- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit

Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.

* fix(vectordb): reject stale index replacements

* fix(vectordb): harden bulk rebuild lifecycle

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-16 18:59:16 +08:00
Qin Haojie d47f2106ee refactor: remove unused and deprecated APIs (#3272)
Delete dead compatibility paths and test-only helpers so unsupported APIs do not remain as accidental contracts.
2026-07-16 10:49:56 +08:00
Yuanqing ZHAOandYuanqing Zhao 7e6a0515f9 perf(cuvs): optimize filters, rebuilds, concurrency, and memory (#3092)
* perf(cuvs): fast-path cached native filter routes

* perf(cuvs): parallelize auto filter preflight

* perf(cuvs): add search route telemetry

* test(cuvs): use a valid telemetry vector dimension

* perf(cuvs): reuse native filter preflight results

* perf(cuvs): allow concurrent snapshot searches

* perf(cuvs): coalesce optional background rebuilds

* perf(cuvs): coordinate per-GPU build admission

* perf(cuvs): add opt-in float16 search

* build(cuvs): support vector benchmark harnesses

* perf(cuvs): bound concurrent GPU searches

* perf(cuvs): avoid partial background rebuilds

* fix(cuvs): address rebuild and telemetry review feedback

* fix(cuvs): defer rebuild until index initialization

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-10 17:22:34 +08:00
Yuanqing ZHAOandYuanqing Zhao 39c778c953 feat: add cuVS vector search backend (#2974)
* feat: add cuVS vector search backend

* docs: add agent memory benchmark strategy

* bench: add cuVS index performance harness

* bench: add public ANN dataset tuning

* docs: record preliminary cuVS index results

* docs: clarify warm index latency

* docs: order cuVS before qdrant

* bench: aggregate independent index runs

* bench: order aggregate variants consistently

* docs: add repeatable index scaling results

* bench: add collection lifecycle benchmark

* docs: add collection lifecycle results

* perf: cache prepared cuvs filters

* docs: report prepared filter cache results

* bench: add async vector concurrency benchmark

* bench: aggregate service concurrency runs

* docs: add async concurrency results

* docs: clarify cuVS dtype behavior

* feat: add memory-aware cuVS auto mode

* feat: reuse native filters for cuVS search

* docs: publish cuVS integration plan as Markdown

* fix: route selective filters before cuVS rebuild

* docs: record selective-first routing results

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-07 12:21:10 +08:00
Jiahui Zhou 21c58bd8f6 Set tags (#2706)
* feat(vectordb): add partial update api

fix(vectordb): support partial updates across adapters

fix(vectordb): return structured update results

test(vectordb): cover update result behavior

* feat(vectordb): support partial upsert semantics

* feat(content): support explicit search tags

* refactor unify search tag update api

* refactor(content): drop unrelated peer_id from semantic refresh

* fix(content): normalize set-tags targets and cli output
2026-06-18 17:48:49 +08:00
eb5b46bed5 fix(vectordb): skip candidates with corrupted JSON fields in search / 检索时跳过 fields 损坏的候选 (#2302)
读取路径容错:search_by_vector 对每条 candidate 的 fields 做防御性 json.loads,
跳过无法解析的损坏数据(例如被存储层 uint16 长度前缀截断的不完整 JSON),并同步
过滤 cands_list / pk_list / scores_list 保持对齐,避免单条坏数据让整批查询抛异常、
进而被上层 index backend 静默降级为空召回。

When a single candidate's `fields` JSON is corrupted, `json.loads` raised
JSONDecodeError and failed the whole batch query; the index backend then
swallowed it and returned [], silently dropping the entire recall. We now skip
the bad candidates while keeping cands_list / pk_list / scores_list aligned,
mirroring the existing None-skip logic right above. Also normalize the index
backend error log to `logger.error(..., exc_info=True)` instead of an inline
traceback dump.

- local_collection.search_by_vector: defensive json.loads + aligned filtering
- viking_vector_index_backend.query: logger.error(..., exc_info=True)
- tests: add regression test test_search_skips_candidate_with_corrupted_fields

Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-29 14:35:53 +08:00
Qin Haojie f1afd40a66 fix(vectordb): reject oversized bytes row strings (#2171) 2026-05-21 19:50:58 +08:00
Qin Haojie 813db21c5a fix(storage): sanitize vectordb unicode recovery (#2103) 2026-05-18 11:58:15 +08:00
MaojiaShengandopenviking ce998873f9 lisence: change the main lisence to AGPL-3.0 (#1085)
* lisence: change the main lisence from Apache-2.0 to AGPL-v3

* lisence: change the main lisence from Apache-2.0 to AGPL-v3

* lisence: change the main lisence from Apache-2.0 to AGPL-v3

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-30 14:37:42 +08:00
Jiahui Zhou 033826f028 Migrate vectordb engine to abi3 packaging (#897)
chore: remove pybind11 remnants after abi3 migration

fix: scope linux abi3 wheel smoke test to engine loader
2026-03-23 19:11:07 +08:00
chenjw 57b320611c FIX: fixes multiple issues in the OpenViking chat functionality and unifies session ID generation logic between Python and Rust CLI implementations. (#446)
* refactor(sandbox): remove docker/aiosandbox backends, simplify SRT config

- Remove docker and aiosandbox from available backends
- Remove settings_path from SrtBackendConfig (now auto-generated in workspace)
- Update SRT settings path to workspace/sandboxes/{session}-srt-settings.json
- Update README examples to use srt backend and remove settingsPath

* refactor(sandbox): remove docker/aiosandbox backends, simplify SRT config

- Remove docker and aiosandbox from available backends
- Remove settings_path from SrtBackendConfig (now auto-generated in workspace)
- Update SRT settings path to workspace/sandboxes/{session}-srt-settings.json
- Update README examples to use srt backend and remove settingsPath

* fix: remove unused handle_chat_direct function and fix unused logs variable

* Fix UTF-8 issues in chat command

* Add tab indentation to Think, Calling, and Result lines in CLI output

* Add first release workflow

* Update release workflow with correct working directory

* 修改 SessionKey 构建逻辑:统一使用 type="cli",channel_id 默认 "default",chat_id 作为 session_id 使用

* Implement machine unique ID as default session ID for ov chat

* Remove unsupported --logs parameter from chat command

* 统一 Python 和 Rust CLI 的默认 session ID 生成逻辑

* 修改日志

* 去掉log依赖

* docs: add VikingBot quick start section to READMEs

* fix: use vikingbot chat instead of ov chat in READMEs

* Revert "fix: use vikingbot chat instead of ov chat in READMEs"

This reverts commit 59f4e87ba0.

* fix: use UUID v4 for machine ID generation in both Rust and Python

* refactor: move truncate_utf8 to utils, fix chat history path, and use BotProcess dataclass

* refactor: update machine ID generation and remove unused chat_v2

- Update Python to use py-machineid library
- Update Rust to use machine-uid crate
- Remove unused chat_v2.rs
- Move machine ID from file storage to system-provided IDs
- Add fallback to "default" if system ID is unavailable

* 优化格式

* ruff format .
2026-03-06 11:12:42 +08:00
Qin HaojieandClaude Opus 4.6 8274cc31ef fix: 修复单测以适配 vectordb 接口重构,统一测试数据路径 (#333)
- 修复 filesystem stat 错误匹配,增加 "not found" 判断
- 为 AddResourceRequest 添加 model_validator 校验 path 参数
- 处理 queue manager 未初始化时 ObserverService 的异常
- 简化 vectordb record ID 生成逻辑,移除 owner_space
- 捕获 HTTP client close 时的 RuntimeError
- 统一测试数据路径至 test_data/ 目录
- 更新测试用例使用 get_context_by_uri() 等新接口
- 移除测试 mock 对 VikingDBInterface 的依赖

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 19:40:12 +08:00
qin-ctx 3165ffa0da feat: add HTTP Server and Python HTTP Client (T2 & T4) (#109)
* feat: add Server/Client architecture with HTTP API and restructure documentation

  - Implement FastAPI-based HTTP server (openviking/server/) with REST API
  - Add client abstraction layer (LocalClient, HTTPClient, BaseClient)
  - Add CLI entry point (python -m openviking serve)
  - Fix bugs: session.session_id, link/unlink param names, hmac.compare_digest
  - Restructure docs: remove numbered prefixes, add guides/, rewrite API reference
    with both Python SDK and HTTP API (curl) examples (en/zh)
  - Add quickstart-server, deployment, authentication, monitoring guides
  - Update examples and design docs to reflect implementation

* 提供单测 和 文档

* Merge branch 'main' into feature/server_client

* feat: add server/client examples and server tests

* fix: cross-references

* fix : tests
2026-02-09 21:12:15 +08:00
kkkwjx c763ab125a refactor: cpp bytes rows (#105)
* refactor: cpp bytes rows

* refactor: cpp bytes row v2
2026-02-09 19:55:26 +08:00
qin-ctx 73642e9508 fix: ignore TestVikingDBProject (#103) 2026-02-09 16:18:39 +08:00
kkkwjx 393a4c5878 Path filter (#92)
* feat: use path field

* refactor: bitmap filter
2026-02-07 17:27:42 +08:00
e4251dc152 Feat:支持原生部署的vikingdb (#84)
* feat: 增加对私部vikingdb的支持

* feat: add search_with_sparse_logit_alpha (#71)

* refactor: Refactor S3 configuration structure and fix Python 3.9 compatibility issues (#73)

* refactor: Refactor S3 configuration structure and fix Python 3.9 compatibility issues
- Extract S3-related configurations from AGFSConfig into a new S3Config class
- Update agfs_manager.py to use the new S3Config structure
- Modify test_config_validation.py to adapt to the new configuration structure
- Update configuration example files to reflect the new configuration structure
- Fix Python 3.9 compatibility issue in embedding_queue.py by replacing `EmbeddingMsg | None` with `Optional[EmbeddingMsg]`

* fix: validate_config validate error

* fix: fix ci (#74)

* refactor: unify async execution utilities into run_async (#75)

- Create centralized run_async() in openviking/utils/async_utils.py
- Remove duplicate _run_async from session.py
- Remove duplicate run_coroutine_sync from observers/async_utils.py
- Replace asyncio.run() calls in sync_client.py with run_async
- Update observers to use the unified run_async utility
- Update tests and documentation for new observer API

* 原生部署的vikingdb由外部来管理

---------

Co-authored-by: kkkwjx <zhoujiahui.01@bytedance.com>
Co-authored-by: baojun-zhang <zhangbaojun.1@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-02-06 23:22:34 +08:00
kkkwjx 9368d838f7 feat: add search_with_sparse_logit_alpha (#71) 2026-02-05 21:28:12 +08:00
kkkwjx 93dee92567 refactor: optimized pyproject.toml with optional groups (#26)
* refactor: optimized pyproject.toml with optional groups

* fix: fix observer
2026-02-02 21:54:01 +08:00
kkkwjx a30cb2e7b9 Lint code (#25)
* feat: pre-commit-config

* style: ruff code
2026-02-02 18:56:15 +08:00
kkkwjx bbbdee90a0 fix: fix agfs-server for windows (#16) 2026-01-31 13:30:06 +08:00
qin-ctx f98dc0ed1c first commit 2026-01-29 20:29:19 +08:00