Commit Graph
485 Commits
Author SHA1 Message Date
Qin HaojieClaude Opus 4.6AutoCoder
2c598e88c5 feat(session): async commit, session metadata, and archive continuity threading (#900)
* feat(session): make commit two-phase with async memory extraction

Session commit now returns immediately after archiving messages (Phase 1).
Summary generation and memory extraction (Phase 2) run in the background
via asyncio.create_task(), returning a task_id for polling progress.

- Add get_task() API across all client layers for querying background task status
- get_session() auto-creates session if it does not exist
- Remove wait parameter and telemetry from commit endpoint
- Add .done completion marker to archive directories
- Update docs (EN/ZH) and tests for new two-phase flow

* feat(session): add .meta.json persistence and auto_create control for get_session

SessionService.get() now defaults to auto_create=False, raising NotFoundError
for missing sessions. A new SessionMeta dataclass tracks created_at, updated_at,
message_count, commit_count, memories_extracted (by category), last_commit_at,
and cumulative llm_token_usage. Meta is persisted to .meta.json and updated on
add_message, commit Phase 1 (message clear), and commit Phase 2 completion
(token usage, memory counts via bind_telemetry). All client layers
(local/async/sync/HTTP) and API docs updated accordingly.

* fix: remove session vectorize

* support commit for openclaw-plugin (#902)

Made-with: Cursor

* fix: reuse latest archive overview in session context

Thread the latest completed archive overview into archive summary generation and memory extraction, and simplify search context assembly to current messages plus the latest archive overview.

Co-Authored-By: Claude Opus 4.6

* refactor: session overview

---------

Co-authored-by: AutoCoder <wulf234@163.com>
2026-03-24 22:30:16 +08:00
chenjwandClaude Opus 4.6 2771765298 Refactor memory extract (#916)
* docs: add memory extractor templating and update mechanism optimization design document

- Add bilingual (English/Chinese) design document for memory templating system
- Include YAML-based MemoryTypeRegistry with 8 built-in types
- Detail ReAct 3+1 phase flow with pre-fetch optimization
- Describe 3-operation Schema: write/edit/delete
- Document RoocodePatch SEARCH/REPLACE format
- Explain dual-mode design: simple mode vs template mode
- Cover pre-fetch optimization: ls directories + read .abstract.md/.overview.md + search once
- Include merge operations: patch, sum, avg, immutable
- Address #578: allow custom prompt template addition and specification

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add memory templating system with ReAct orchestrator

- Add YAML-configurable memory schemas (cards, events, entities, etc.)
- Implement MemoryReAct with tool use (read/find/ls)
- Add schema-driven memory operations (write_uris/edit_uris/delete_uris)
- Implement memory patch handler for incremental updates
- Add comprehensive test suite

* refactor: memory extractor templating system with ReAct orchestrator

## Summary

Implement memory templating system (GitHub Issue #578) - a complete
rewrite of the memory extractor subsystem to support YAML-configurable
memory types instead of hardcoded categories.

## Key Changes

### Architecture
- Replace hardcoded 8 memory types with YAML-configurable schema system
- Add MemoryTypeRegistry to load memory type definitions from YAML files
- Dynamic Pydantic model generation from schema for type safety
- Field-level merge operations: PATCH, SUM, IMMUTABLE

### Memory Extraction Flow
- Implement ReAct orchestrator for single-pass memory updates
- MemoryUpdater for applying operations to storage
- Memory tools (read, search, ls) for ReAct loop
- Stable JSON parser with 5-layer fault tolerance

### File Naming & Storage
- Semantic filenames from template ({topic}.md instead of random IDs)
- Two memory modes: simple mode and template mode
- MEMORY_FIELDS HTML comment for structured metadata

### Configuration
- 9 YAML templates in openviking/prompts/templates/memory/
- memory_config.py for memory system configuration
- Dual-threshold compact upload mechanism in design doc

### Deletions
- Remove old memory_content.py, memory_data.py, memory_operations.py
- Remove memory_types.py, memory_utils.py, memory_patch.py
- Remove corresponding old test files

### Updated Components
- VLM backends (litellm, openai, volcengine) for new interfaces
- Session and service core integration
- Test suite for new architecture

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: pass ctx/user/session_id in commit_async for memory extraction

## Summary

Fix missing parameters in commit_async() when calling extract_long_term_memories().
The synchronous commit() method correctly passes these parameters, but the async
version was missing them, causing memory extraction to be skipped.

## Changes

- Pass user=self.user, session_id=self.session_id, ctx=self.ctx
  in commit_async() when calling extract_long_term_memories()

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: convert FindResult to dict before returning from search tool

## Summary

Fix JSON serialization error by converting FindResult object to dict
using its to_dict() method before returning from MemorySearchTool.

## Changes

- In MemorySearchTool.execute(), return search_result.to_dict()
  instead of search_result directly

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: swap None check before accessing final_operations in memory_react

Also rename schema_models.py to schema_model_generator.py for clarity.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add edit_overview support and optimize memory registry initialization

- Add edit_overview_operations to MemoryUpdater for updating .overview.md files
- Optimize MemoryTypeRegistry initialization in SessionCompressorV2 (load once)
- Various memory templating system improvements

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: remove unnecessary indent in JSON schema output

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: add markdown link format hint to overview field description

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add pre-fetch search based on user messages in conversation

Also fix duplicate line in system prompt.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* rebase

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-24 15:28:46 +08:00
Yaoyao 8ab471d8a6 add latest wechat-group-qrcode (#919)
the existed one has been out of date.
2026-03-24 14:36:21 +08:00
352cd89c0e feat(retrieve): add provenance metadata to search results (#852)
* feat(retrieve): add provenance metadata to search results

Adds an opt-in `include_provenance` parameter to the search/find API
endpoints. When enabled, the response includes a `provenance` array
with per-query retrieval details: which directories were traversed,
which tier (L0/L1/L2) each result came from, match reasons, and the
full thinking trace.

The internal data was already being collected in MatchedContext.level,
MatchedContext.context_type, and QueryResult.thinking_trace. This
change surfaces it through the API for retrieval observability, which
the README lists as a core design goal ("Visualized Retrieval
Trajectory").

Backward compatible: defaults to false, existing clients see no change.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: add provenance feature screenshot

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-24 11:14:37 +08:00
r266-techandwzr 2c4137346e docs: add Feishu/Lark cloud document parser documentation (#906)
Add documentation for the Feishu/Lark cloud document parser (PR #831):

- Add Feishu to supported formats table in resources docs (EN + ZH)
- Add Feishu URL usage examples with Python SDK, HTTP API, and CLI
- Document feishu config section with all parameters (EN + ZH)
- Include setup instructions, dependency info, and Lark international note
- Cover all 4 supported document types: docx, wiki, sheets, bitable

Closes #859

Co-authored-by: wzr <2668940489@qq.com>
2026-03-24 11:08:40 +08:00
50e1ff91d2 feat(cli): add ov doctor diagnostic command (#851)
* feat(cli): add ov doctor diagnostic command

Adds a new `ov doctor` command that validates all OpenViking subsystems
and reports actionable diagnostics without requiring a running server.

Checks: config file, Python version, native vector engine (PersistStore),
AGFS, embedding provider, VLM provider, and disk space. Each check is
isolated so one failure doesn't block others, and every failure includes
a specific fix suggestion.

This addresses a real pain point: when the native engine is missing from
a pip wheel (e.g., Python 3.13), the only feedback is 50+ ERROR log
lines with no actionable guidance. `ov doctor` catches this immediately:

  Native Engine: FAIL  No compatible engine variant
    Fix: pip install openviking --upgrade --force-reinstall
    Alt: Use vectordb.backend = "volcengine" instead of "local"

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: add ov doctor screenshot examples

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(doctor): reuse resolve_config_path and drop hardcoded OPENAI_API_KEY fallback

Address review feedback from @qin-ctx:
- Replace _CONFIG_SEARCH_PATHS and _find_config() with resolve_config_path()
  from config_loader.py to avoid two sources of truth for config discovery
- Remove hardcoded OPENAI_API_KEY env var fallback from embedding and VLM
  checks - only check the api_key field in config
- Renumber inline comments in rust_cli.py (1/2/3 matching docstring)

---------

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-24 00:34:57 +08:00
baojun-zhang 5bbb22a6ef feat: add encrypt doc && refactoring encrypt code (#893)
* feat: add encrypt doc && fix Partial reads/wrong location problem && refactoring duplicated code && add

* feat: reformat code

* feat: resolve comment

* feat: resolve comment

* feat: resolve comment
2026-03-23 19:23:11 +08:00
3b5b9a42c0 feat(openclaw-plugin):context engine refactor design & enforce token budget and reduce context bloat (#891)
* fix(openclaw-plugin): enforce token budget and reduce context bloat (#730) (#796)

* test(openclaw-plugin): add vitest test infrastructure for #730

* fix(openclaw-plugin): raise recallScoreThreshold default from 0.01 to 0.15 (#730)

* fix(openclaw-plugin): narrow isLeafLikeMemory boost to level-2 only (#730)

* fix(openclaw-plugin): prefer abstract over full content fetch in memory injection (#730)

* feat(openclaw-plugin): add recallMaxContentChars and recallPreferAbstract config (#730)

* feat(openclaw-plugin): enforce tokenBudget in injection with decrement loop (#730)

* fix(openclaw-plugin): update recallScoreThreshold placeholder to match new default (#730)

* fix(openclaw-plugin): deduplicate content resolution and document budget behavior (#730)

Extract resolveMemoryContent() helper to eliminate duplicate content-resolution
logic between buildMemoryLines and buildMemoryLinesWithBudget. Add JSDoc and
inline comment documenting intentional first-line budget overshoot (spec §6.2).
Tighten test assertion from <=120 to <=106 tokens.

* fix(openclaw-plugin): use truthy fallback for empty abstract strings (#730)

Change nullish coalescing (??) to truthy fallback (||) in
resolveMemoryContent() so empty-string abstracts fall back to item.uri
instead of producing empty content lines.

* docs: add openclaw context engine refactor design

Co-authored-by: Mijamind <mijamind@163.com>
Co-authored-by: GPT-5.4 <noreply@openai.com>

* add afterTurn refactor in design_doc

* add afterTurn compact in design_doc

---------

Co-authored-by: chethanuk <chethanuk@outlook.com>
Co-authored-by: GPT-5.4 <noreply@openai.com>
Co-authored-by: wlff123 <wulf234@163.com>
Co-authored-by: xuwengui <huangxun375@gmail.com>
2026-03-23 15:40:15 +08:00
yepper c1480dd500 docs(api): add documentation for incremental update feature (#886)
Explain the incremental update behavior when calling `add_resource()` repeatedly for the same URI. Describe the trigger conditions, semantic stage optimizations, and filesystem/index synchronization process.
2026-03-23 13:13:06 +08:00
chethanuk 5633394583 docs: add Gemini embedding provider usage and installation guide (#830) 2026-03-22 10:03:01 +08:00
ChenXiaofeiandClaude Sonnet 4.6 44d9542fc3 feat(rerank): add OpenAI-compatible rerank provider (#785)
- Add OpenAIRerankClient using standard flat request/response format
  compatible with DashScope compatible-api and other OpenAI/Cohere-style
  rerank APIs (no input/output wrappers)
- Fix silent data corruption: add index bounds-checking so out-of-bounds
  or missing index returns None with a warning
- Add provider allow-list validation in RerankConfig ('vikingdb'|'openai')
- Remove unnecessary getattr() in RerankClient.from_config()
- Update ov.conf.example: keep vikingdb (doubao) as primary rerank config,
  add rerank_openai_example section for DashScope qwen3-rerank
- Update docs (en/zh): add OpenAI-compatible provider example alongside
  existing volcengine example in configuration guide and schema
- Add 24 tests covering success, edge cases, and factory dispatch

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-20 14:22:25 +08:00
Qin Haojie 7be9514ed7 fix(observer): pass RequestContext through vikingdb observer for tenant-scoped vector count (#807)
VikingDBObserver.count() was called without RequestContext, falling back
to account_id="default" and filtering out vectors from other tenants.
Thread ctx from the HTTP router through ObserverService and VikingDBObserver
so each tenant only sees their own vector count.

Closes #786

Credit-To: jackjin1997
2026-03-20 13:47:32 +08:00
Zayn Jarvis 20ec7e656e feat: update readme logo (#799) 2026-03-20 11:55:00 +08:00
Xiaogang Zhouandxiaogang.zhou 1770a1792a feat(embedder): minimax embeding (#624)
* feat: enable minimax embedding and adapt to separate query/document parameter configuration

* feat: enable minimax embedding and adapt to separate query/document parameter configuration

---------

Co-authored-by: xiaogang.zhou <xiaogang.zhou@bytedance.com>
2026-03-20 00:36:51 +08:00
dingbenanddingben.db@bytedance.com <dingben.db@bytedance.com@bytedance.com> ff751a4304 docs: add docker compose instructions and mac port forwarding tip to quickstart (#781)
Co-authored-by: dingben.db@bytedance.com <dingben.db@bytedance.com@bytedance.com>
2026-03-19 20:58:32 +08:00
Qin HaojieandClaude Opus 4.6 b31c477fb4 feat(client): add account/user params for root key multi-tenant auth (#767)
SDK HTTP client now supports account and user parameters so root key
holders can access tenant-scoped data APIs without being rejected by
the server explicit-tenant requirement.

Co-Authored-By: Claude Opus 4.6
2026-03-19 16:04:13 +08:00
zhoujiahui f467de4ed9 feat: add memory extract telemetry breakdown (#735)
feat: add resource telemetry breakdown

telemetry: omit zero-valued summary fields

feat(resources): add temp upload telemetry support

docs: move telemetry guide out of design

docs: sync contributor build requirements
2026-03-19 15:10:59 +08:00
KorenKrita d739a5be35 feat(vlm): add streaming response handling for OpenAI VLM (#756)
* feat(vlm): add streaming response handling for OpenAI VLM

Add configurable stream mode for OpenAI-compatible VLM providers.
This enables proper handling of SSE (Server-Sent Events) responses
from providers that force streaming format.

Changes:
- VLMBase extracts stream from config (default: False)
- OpenAIVLM adds _extract_from_chunk(), _process_streaming_response(),
  _process_streaming_response_async() for streaming response processing
- All API methods (get_completion, get_completion_async, get_vision_completion,
  get_vision_completion_async) now support stream parameter
- VLMConfig adds stream field with migration to providers structure
- Update configuration docs (zh/en) with stream parameter documentation
- Add comprehensive tests for stream configuration

Co-Authored-By: KorenKrita <KorenKrita@gmail.com>

* cr fix: fix MagicMock async generator issue in VLM tests
2026-03-19 11:40:31 +08:00
KorenKrita cd87c0aaa6 Add support for custom HTTP headers in VLM models (OpenAI-compatible) (#723)
* feat(vlm): add custom HTTP headers support for OpenAI-compatible backends

Add extra_headers configuration option for VLM models to support custom
HTTP headers (e.g., HTTP-Referer, X-Title) when using OpenAI-compatible
providers like OpenRouter.

Changes:
- VLMBase extracts extra_headers from config
- OpenAIVLM passes extra_headers as default_headers to OpenAI client
- VLMConfig supports extra_headers in providers config
- Add tests for extra_headers functionality
- Update configuration docs (zh/en) and example config

Co-Authored-By: KorenKrita <KorenKrita@gmail.com>

* style: fix ruff formatting issues in test file

* cr fix: add extra_headers field to VLMConfig and clean up example

- Add extra_headers: Optional[Dict[str, str]] field to VLMConfig
- Migrate extra_headers to providers structure in _migrate_legacy_config
- Remove confusing example-only keys from ov.conf.example
- Add test for flat extra_headers config style

Co-Authored-By: KorenKrita <KorenKrita@gmail.com>

* style: fix ruff formatting in test file

Format long assertion lines to pass CI checks.

Co-Authored-By: KorenKrita <KorenKrita@gmail.com>
2026-03-18 16:31:12 +08:00
yepper 834b808183 feat(resources): add resource watch scheduling and status tracking (#709)
* feat(resource): add watch interval support for resource monitoring

implement resource watch functionality that allows automatic monitoring and re-processing of resources at specified intervals. key features include:
- add watch_interval parameter to resource APIs
- create watch scheduler service for task execution
- handle conflict detection for active watch tasks
- provide watch status query capability
- include comprehensive tests and examples

the watch feature enables periodic automatic updates of resources without manual intervention, improving data freshness for frequently changing content

* feat(resources): add watch status tracking and improve resource processing

- Implement get_watch_status API for tracking resource watch status
- Add immediate persistence for first-time resource additions
- Improve file change detection with size comparison
- Refactor watch scheduler with better concurrency control
- Add test coverage for watch status and resource processing
- Remove unused watch manager references and clean up code

* refactor: improve code style and fix minor issues

- Simplify logging by removing redundant data copying
- Fix syntax errors in docstrings and string literals
- Add new fields to EmbeddingMsg class
- Improve line wrapping and formatting
- Update watch task storage URIs to use hidden files

* refactor(embedding_msg): simplify EmbeddingMsg constructor by removing unused fields

Remove media_uri, media_mime_type and id parameters as they are not used in the implementation

* feat(watch): add backup task recovery and simplify permission check

Add test case for recovering tasks from backup storage when primary is missing
Remove require_owner parameter from _check_permission as it's redundant with the existing role-based checks

* feat(resources): add watch_interval support for resource updates

Add watch_interval parameter to enable periodic resource updates. When target is specified, watch_interval > 0 creates/updates a watch task, while <= 0 disables it. Also simplify resource moving logic in ResourceProcessor by using direct mv operation.

* refactor(watch): remove deprecated get_watch_status functionality

remove get_watch_status method and related tests, update examples to use direct task access
update watch manager to use ConflictError for URI conflicts and include original_role in tasks
add validation for watch_interval requiring target URI

* fix(resource_service): validate watch interval before processing resource

Move watch interval validation earlier in the flow to fail fast when 'to' parameter is missing

* refactor(resource_processor): remove redundant temp_uri assignment
2026-03-18 15:16:14 +08:00
Qin HaojieandClaude Opus 4.6 1823a7c4f7 feat(storage): add path locking and selective crash recovery for write operations (#431)
* feat(storage): add transaction support with journal, undo, and crash recovery

Implement a full transaction system for VikingFS storage operations including
write-ahead journal, path locking, undo/rollback, context manager API, and
crash recovery. Includes comprehensive tests and documentation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test(transaction): add e2e rollback tests for mv and multi-step operations

Add end-to-end tests covering rollback scenarios that were missing:
- mv rollback: file moved back to original location on failure
- mv commit: file persists at new location
- Multi-step rollback: mkdir + write + mkdir all reversed in order
- Partial step rollback: only completed entries are reversed
- Nested directory rollback: child removed before parent
- Best-effort rollback: single step failure does not block others

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(storage): add transaction support with path locking and journal

Implement transaction system for VikingFS with ACID-like guarantees:
- TransactionManager with configurable lock timeout and journal-based recovery
- PathLock supporting point, subtree, and mv lock modes
- Refactor VikingFS mv to use cp+rm to prevent lock files from being carried
- Fix stale lock detection returning false for missing lock files
- Update ragas eval to use LangchainLLMWrapper

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: tests

* fix(transaction): fix rollback and race condition bugs

- Reconstruct RequestContext from undo params for vectordb_delete/update_uri
  rollback (previously skipped silently due to missing ctx)
- Serialize ctx fields into undo params in rm/mv operations
- Fix Phase 1 undo path to target archive dir instead of session root
- Remove Phase 2 fs_write_new undo (overwrites are idempotent, checkpoint
  handles recovery)
- Add ancestor SUBTREE recheck after lock creation in acquire_subtree
- Move _collect_uris inside TransactionContext in rm/mv to close race window
- Log journal persistence failures instead of silently swallowing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(transaction): make TransactionManager required and rewrite tests with real backends

Remove all optional/fallback code paths where tx_manager could be None. get_transaction_manager()
now raises RuntimeError if not initialized. Fix undo rollback to reconstruct ctx for vectordb_upsert
and use correct agent_id default. Replace mock-based transaction tests with integration tests using
real AGFS and VectorDB backends.

* refactor(transaction): make rollback fully async and unify session commit path

- Convert execute_rollback/rollback_entry to async, removing sync run_async wrappers
- Unify Session.commit() to delegate to commit_async(), removing duplicate phase methods
- Fix SUBTREE lock to conflict with ancestor SUBTREE locks (was previously missing)
- Fix mv lock mode: directory moves now use SUBTREE on both source and destination
- Replace deprecated asyncio.get_event_loop() with get_running_loop()
- Remove max_parallel_locks config option
- Update docs (en/zh) and tests to match new async rollback signatures

* fix: tests

* refactor(transaction): simplify session commit and add redo-based crash recovery

Session commit no longer wraps archive phase in a transaction. Phase 2 uses
redo semantics so crashed memory-extraction can be replayed from archive.
PathLock stale-lock cleanup no longer redundantly re-checks timeout.
Semantic processor vectorization runs concurrently via asyncio.gather.

* fix: transaction

* fix: UserIdentifier

* refactor(transaction): replace undo-based transaction manager with lightweight lock + redo-log

Remove the heavyweight TransactionManager/Journal/UndoEntry system (~4000 lines) and
replace it with a simpler architecture: LockManager for path locking, LockContext as
the async context manager, LockHandle/LockOwner protocol, and a RedoLog for crash
recovery of session_memory operations. VikingFS rm/mv now use inline error handling
instead of rollback semantics. Updated docs, observers, and tests accordingly.

Co-Authored-By: Claude Opus 4.6

* fix(transaction): remove checkpoint dead code, fix TOCTOU race, clarify mv lock param

- Remove unused _write_checkpoint/_write_checkpoint_async/_read_checkpoint
  from Session (superseded by redo-log)
- Re-resolve URI inside lock in resource_processor Phase 3.5 to prevent
  concurrent add_resource calls from resolving to the same final_uri
- Rename acquire_mv dst_path to dst_parent_path with docstring to clarify
  that callers pass the destination parent directory

* fix: path

* fix: resource lock

* fix: test

* docs: update

* fix: tests

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-18 14:50:19 +08:00
Qin Haojie ae35f46c75 Revert "feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF) (#607)" (#703)
This reverts commit 95bd19797f.
2026-03-17 21:29:21 +08:00
chethanuk 95bd19797f feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF) (#607)
* feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF)

Native text + multimodal (image, video, audio, PDF) embedding via `gemini-embedding-2-preview` (google-genai 1.67.0). Additive provider pattern — Volcengine remains the default; Gemini is opt-in via `provider: "gemini"` in `ov.conf`.
    - **Model**: `gemini-embedding-2-preview`
    - **Input**: text, image, video, audio, PDF (17 MIME types)
    - **Output dimension**: 128–3072 (default: **3072**, recommended: 768 / 1536 / 3072)
    - **Input token limit**: **8,192 tokens**
    - **Supported MIME types**: `image/jpeg`, `image/png`, `image/gif`, `image/webp`, `audio/mpeg`, `audio/mp3`, `audio/wav`, `audio/ogg`, `audio/flac`, `video/mp4`, `video/mpeg`, `video/mov`, `video/avi`, `video/webm`, `video/wmv`, `video/3gpp`, `application/pdf`
- Gemini Embedding 2 Multimodal Support: Introduced a new GeminiDenseEmbedder to support native text and multimodal (image, video, audio, PDF) embedding using the gemini-embedding-2-preview model. This is an opt-in provider via configuration.
- Extended Queue Pipeline for Multimodal Content: The EmbeddingMsg now carries media_uri and media_mime_type to facilitate multimodal content processing. The TextEmbeddingHandler.on_dequeue() method was updated to read raw bytes from viking_fs and call embed_multimodal() when applicable.
- End-to-End Configuration and Security: The EmbeddingConfig now registers the 'gemini' provider with a task_type field. A critical security validation was added to ensure media_uri matches context_data['uri'] before file reads, preventing forged queue messages from accessing arbitrary files. If validation fails or multimodal embedding fails, it falls back to text embedding.
- Multimodal Content Representation: A new ModalContent dataclass was introduced to represent media references, including MIME type, URI, and optional raw data, enabling the Vectorize object to encapsulate both text and media for embedding

* feat: Add asynchronous batch embedding with concurrency control to Gemini embedder.

* Reduce scope to use GeminiDenseEmbedder as only text embed
2026-03-17 21:15:12 +08:00
Qin HaojieandClaude Opus 4.6 ded21cb432 feat(session): add used() endpoint and overhaul plugin install docs (#684)
Add POST /sessions/{session_id}/used API to record actually used
contexts and skills, enabling active_count tracking on commit.
Overhaul all plugin installation docs: simplify to npm global install,
add OpenClaw >= 2026.3.12 compatibility warning, expand troubleshooting
and configuration reference.

Co-Authored-By: Claude Opus 4.6
2026-03-17 14:17:50 +08:00
kfiramar cdd0f795c3 feat(embedding): add Voyage dense embedding support (#635)
* feat(embedding): add voyage dense embedder

Add first-class Voyage dense embedding support with a dedicated embedder,
provider validation, model-aware default dimensions, and focused tests.

Keep the configuration surface intentionally narrow:
- use the existing dimension field and map it to Voyage's
  output_dimension request field
- do not expose Voyage-only output_dtype
- do not expose query/document mode until OpenViking has separate
  index/query embedder configuration

This keeps the PR aligned with the current OpenViking architecture,
which stores and retrieves dense float vectors through a single dense
embedder configuration.

Verification:
- .venv/bin/python -m pytest tests/unit/test_voyage_embedder.py tests/unit/test_embedding_config_voyage.py --noconftest -o addopts='' -q
- .venv/bin/python -m pytest tests/misc/test_config_validation.py -o addopts='' -q
- .venv/bin/ruff check openviking/models/embedder/voyage_embedders.py openviking_cli/utils/config/embedding_config.py tests/unit/test_embedding_config_voyage.py tests/unit/test_voyage_embedder.py

* refactor(embedder): inline Voyage extra body

Remove the trivial Voyage-specific helper and inline the output_dimension payload construction at the two call sites.

This keeps the request shape unchanged while making the embedder implementation more direct.
2026-03-16 21:46:37 +08:00
ChenXiaofei e5abe88f04 feat(embedding): add Ollama provider support for local embedding (#644)
- Add 'ollama' as a supported embedding provider

- Ollama runs locally via OpenAI-compatible API, no API key required

- Allow OpenAI provider to work without api_key when api_base is set

  (supports local OpenAI-compatible servers like vLLM, LocalAI)

- Add configuration example and tests for Ollama provider

This enables fully local embedding deployment without cloud API keys.
2026-03-16 17:55:50 +08:00
Yaoyao 18f120671a Add new wechat group qrcode image (#649)
wechat group6  qrcode nupdated
2026-03-16 12:01:40 +08:00
zhoujiahui b280b56b30 feat(trace): add request-level trace metrics and API support (#640)
refactor: replace operation trace with telemetry

fix telemetry demo skill ingestion

simplify telemetry summary metric keys

rename remaining trace telemetry artifacts

feat: support configurable telemetry payloads

docs: rewrite operation telemetry design in chinese

fix: reject telemetry for async session commit

refactor: isolate telemetry orchestration

refactor: remove telemetry from find payloads

refactor: remove telemetry event payloads

fix(trace): keep only telemetry-related changes

fix(trace): remove top-level usage from telemetry responses

feat(console): default telemetry on proxied operations
2026-03-15 22:44:15 +08:00
mildred522 73c7072429 fix: integrate rerank into hierarchical retriever (#599)
* feat(retrieval): integrate rerank into hierarchical retriever

* chore(retrieval): align retriever typing with lint checks

* docs(retrieval): clarify rerank fallback in zh docs
2026-03-14 17:46:41 +08:00
chenjw 4cf688852b feat: add --sender parameter to chat commands (#562)
* feat: add --sender parameter to chat commands

- Add --sender option to Python CLI chat command
- Add --sender option to Rust CLI chat command
- Pass sender ID through to channels
- Display sender in interactive mode header
- Update langfuse integration for compatibility

* fix: align default sender ID to "user" for consistency

Align default sender ID in SingleTurnChannel from "default" to "user" to
match ChatChannel's default, ensuring consistent sender identification
across interactive and single-turn chat modes.

* style: add trailing comma for consistency

* fix: align Rust CLI default sender to "user"

Change Rust CLI's default sender from "cli_user" to "user" to match
Python side (ChatChannel and SingleTurnChannel), ensuring consistent
default sender identification across both Rust and Python CLI tools.

* fix: pass through total_tokens in langfuse usage conversion

When converting from old usage format (prompt_tokens/completion_tokens/total_tokens)
to usage_details format, also pass through total_tokens as 'total' field if
it's available in the usage dict.

* fix: protect langfuse.flush() with try/except

Wrap self.langfuse.flush() calls in try/except blocks to prevent
flush failures from discarding successfully obtained LLM responses.

- In success path: flush() failure only logs debug message
- In error path: flush() failure silently ignored (already in error handling)

* fix: disable langfuse propagate_attributes to fix generator error

Temporarily disable langfuse propagate_attributes context manager to fix
RuntimeError: generator didn't stop after throw(). The context manager
had exception handling issues when exceptions were thrown inside the block.

This preserves the API while avoiding the runtime error.

* fix: properly implement langfuse propagate_attributes without generator error

Reimplement propagate_attributes with manual __enter__/__exit__ management to
avoid the 'generator didn't stop after throw()' error. Key changes:

- Use local variable to avoid name shadowing with the method
- Only catch exceptions when entering the context manager
- Let inner block exceptions propagate normally
- Always exit the context manager in finally block
- Restore session_id/user_id propagation to langfuse
2026-03-13 14:02:05 +08:00
yangxinxin-7andClaude Sonnet 4.6 e46eeaf47f fix: correct Volcengine sparse/hybrid embedder and update sparse model docs (#561)
* fix: correct Volcengine sparse/hybrid embedder and update sparse model docs

修复火山引擎 Sparse/Hybrid Embedder 并更新 Sparse 模型文档

**Bug fixes / 问题修复:**
- Fix `Ark()` init crash when `api_base` is None by only passing `base_url` when set
  修复 `api_base` 为 None 时 `Ark()` 初始化崩溃问题,仅在有值时传入 `base_url`
- Fix `VolcengineSparseEmbedder.embed()` using `response.data[0]` — the multimodal
  API always returns a single object, not a list
  修复 `embed()` 错误使用 `response.data[0]`,multimodal API 始终返回单个对象而非列表
- Fix `embed_batch()` for sparse and hybrid embedders: the multimodal API input array
  is for multi-modal inputs of a single sample, not batching; delegate to `embed()` per text
  修复 Sparse/Hybrid 的 `embed_batch()`:multimodal API 的 input 数组是单样本多模态输入,
  不支持 batch,改为逐条调用 `embed()`

**Docs / 文档:**
- Update sparse model from `bm25-sparse-v1` to `doubao-embedding-vision-250615` in EN/ZH docs
  中英文文档中 sparse 模型由 `bm25-sparse-v1` 更新为 `doubao-embedding-vision-250615`
- Add note that Volcengine sparse embedding is supported from `doubao-embedding-vision-250615`
  and only supports text input
  新增说明:火山引擎 Sparse embedding 从 `doubao-embedding-vision-250615` 起支持,仅支持文本输入

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs: remove text-only restriction note from sparse embedding docs

文档:移除 Sparse embedding 中仅支持文本输入的限制说明

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-12 21:18:36 +08:00
Yaoyao 816903bdaf Add files via upload (#543) 2026-03-12 14:36:44 +08:00
r266-techZaynJarvis (review feedback)r266-tech
277d6e21df docs: add MCP integration guide (EN + ZH) (#518)
Add comprehensive MCP integration documentation covering:
- Transport mode comparison (HTTP vs stdio)
- Multi-session safety warnings for stdio contention
- Client configuration for Claude Code, Claude Desktop, Cursor, OpenClaw
- Available MCP tools reference
- Troubleshooting common errors

Addresses documentation improvement requested in #473.

Co-authored-by: ZaynJarvis (review feedback)

Co-authored-by: r266-tech <r266-tech@users.noreply.github.com>
2026-03-11 22:43:10 +08:00
chuanbao666 8c740fd565 chore: 编译子命令失败报错, golang版本最低要求1.22+ (#444)
* chore: 编译子命令失败报错, golang版本最低要求1.22+

* fix: agfs默认启用binding-client相关改造
2026-03-05 20:52:38 +08:00
chuanbao666 c1537a5dc5 chore: agfs-client默认切到binding-client模式 (#430) 2026-03-05 15:45:05 +08:00
MaojiaShengandopenviking c4209da6cc chore: downgrade golang version limit, update vlm version to seed 2.0 (#425)
* feat: define a system path for future deployment

* feat: define a system path for future deployment

* fix: golang downgrade to 1.19

* fix: golang downgrade to 1.19, and change doc

* docs: change model recommendation

* fix: golang downgrade to 1.19, and change doc

* fix: mv test files

* fix: loguru

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-05 14:34:02 +08:00
MaojiaSheng 079efd1177 docs: pip upgrade suggestion (#418) 2026-03-04 20:48:32 +08:00
MaojiaShengandopenviking ab849f44f2 docs: use openviking-server to launch server (#398)
* feat: support glob -n and cmd echo

* feat: server launch commands

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-03 18:37:05 +08:00
MaojiaShengandopenviking e4b1ffed10 Feat: CLI optimization (#389)
* feat: remove python cli and disable python -m openviking

* feat: ls, tree, find, search, grep all use --node-limit as the result limiting arg

* feat: add ls -n

* feat: update agfs to support grep -n

* feat: update agfs to support grep -n

* docs: cancel modify

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-03 13:23:37 +08:00
Qin Haojie a75b17f305 docs/update_wechat (#377) 2026-03-02 16:46:06 +08:00
yangxinxin-7 c02b4e8b1a eat(opencode): add opencode plugin and update docs (#351) 2026-02-28 17:58:56 +08:00
baojun-zhang 38a29c09ea docs: add storage configuration guide (#329)
* docs: add storage configuration guide

* docs: add storage configuration guide
2026-02-27 18:51:19 +08:00
Qin HaojieandClaude Opus 4.6 048c201b73 feat(docs,ci): 完善云上部署文档与 Docker CI 流程 (#320)
- release workflow 新增 Docker 镜像构建推送至 GHCR
- 重写 examples/cloud/GUIDE.md 云上部署指南,补充架构概览、systemd/Docker/Helm 三种部署方式、多租户管理、FAQ 等章节
- ov.conf.example 占位符改为 <your-xxx> 格式,移除 rerank 配置
- 中英文部署文档(docs/zh|en/guides/03-deployment.md)对齐:Docker 改用 GHCR 预构建镜像,Kubernetes 改用 Helm chart,并引用 GUIDE.md

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 08:47:22 +08:00
chuanbao666 4c9340fc10 feat(agfs): agfs新增binding client (#304)
* feat(agfs): agfs新增binding client

fix: viking_fs适配binding client, 简化agfs相关参数

* feat(agfs): agfs binding-client支持windows

* docs: 补充agfs binding-client用法

* fix(agfs): 默认带上agfs binding-client依赖库; setup.py支持编译该库

* fix(agfs): add agfs binding-client linux so

* fix(agfs): code format
2026-02-26 20:27:00 +08:00
Qin HaojieandClaude Opus 4.6 9e69113ad0 fix(server): 未配置 root_api_key 时仅允许 localhost 绑定 (#310)
当 root_api_key 未配置时 resolve_identity() 将所有请求解析为 ROOT,
结合默认绑定 0.0.0.0 会导致任何网络请求均可执行管理员操作。

- 将默认 host 从 0.0.0.0 改为 127.0.0.1
- 添加 validate_server_config() 启动校验:无 key + 非 localhost 时拒绝启动
- 将 dev mode 日志从 info 升级为 warning
- 更新中英文认证文档的开发模式段落

Closes #302

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 17:58:24 +08:00
Qin HaojieandClaude Opus 4.6 1a40839466 feat(client): support timeout config in ovcli.conf and fix Rust CLI alignment (#308)
- Python HTTP client reads `timeout` from ovcli.conf when using default value,
  with priority: SDK explicit param > ovcli.conf > default 60.0
- Rust CLI: replace unused `user` field with `agent_id`, send X-OpenViking-Agent
  header to align with Python client behavior
- Rust CLI: read `timeout` from ovcli.conf and pass to reqwest client
- Move timeout documentation from configuration guide to API overview
- Update examples and ovcli.conf.example with timeout field

Closes #306

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 17:25:42 +08:00
yepper 155866ce52 对齐 docs 和代码里 storage config 的 workspace 属性 (#289)
* 对齐 docs 和代码里 storage config 的 workspace 属性

* workspace 保持统一 ./data

* docs: update configuration docs with new parameters

- Add `max_concurrent` parameter for embedding configuration
- Rename `base_url` to `api_base` in VLM configuration
- Add `thinking` and `max_concurrent` parameters for VLM
- Update AGFS backend options and timeout value
2026-02-25 22:24:05 +08:00
kkkwjxandqin-ctx 37b5d0d359 Multi tenant (#283)
* test: add multi-tenant isolation unit coverage

* feat: implement phase2 multi-tenant context isolation

* feat: multi tenant

* fix: resolve multi-tenant merge conflicts

* feat: multi tenant

* fix: enforce tenant isolation for session get

* tos config

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-02-25 18:25:36 +08:00
MaojiaShengandopenviking a412c7d164 feat: break change, remove is_leaf scalar and use level instead (#271)
* feat: wrap log configs in LogConfig, add log rotation, fix empty file creation
- Create new LogConfig class to wrap all log-related configuration
- Update OpenVikingConfig to use nested log config instead of separate fields
- Add log rotation support with TimedRotatingFileHandler in logger.py
- Add configure_uvicorn_logging to make uvicorn use OpenViking's logging config
- Fix empty file being created in current directory by properly handling log.output=file in logging_init.py
- Update examples/ov.conf.example to use new nested log config structure
- Update __init__.py to export LogConfig and initialize_openviking_config

* docs: help to configure log and workspace

* feat: break change, remove is_leaf scalar and use level instead

* feat: break change, remove is_leaf scalar and use level instead

* feat: rust cli add-resource support zip

---------

Co-authored-by: openviking <openviking@example.com>
2026-02-25 13:01:41 +08:00
SeanZ ea93ad04e9 Feat/add parts support to http api (#270)
* feat(session): add parts support to HTTP API add_message endpoint

- Add optional 'parts' parameter to AddMessageRequest
- Support two modes: simple (content string) and parts (array)
- Update LocalClient and HTTP Client for consistency
- Update Chinese and English documentation

This enables HTTP API clients to store full Part information
(TextPart, ContextPart, ToolPart) instead of just text content,
achieving feature parity with the Python SDK.

* fix(storage): wrap AGFSClientError as FileNotFoundError in read_file

When reading a non-existent file, agfs.read() raises AGFSClientError
which was not being caught by Session.load()'s exception handler.
This caused the server to return 422 errors instead of gracefully
handling missing session files.

Now read_file() catches all exceptions from agfs.read() and re-raises
them as FileNotFoundError, ensuring consistent error handling across
the codebase.

* fix(storage): distinguish error types in VikingFS read operations

Map AGFSClientError to appropriate Python exceptions (FileNotFoundError,
PermissionError, IOError) instead of always raising FileNotFoundError.
This allows callers to handle file-not-found vs network errors differently.

* fix(client): restore read params and refactor part conversion

- Restore offset/limit parameters in HTTPClient and LocalClient read()
  methods that were accidentally removed
- Extract part_from_dict() to openviking/message/part.py to reduce
  code duplication between LocalClient and sessions router
2026-02-25 12:42:18 +08:00