Commit Graph
69 Commits
Author SHA1 Message Date
chenjwandqin-ctx ed1bd4b897 refactor(memory): 统一 V3 提取并提升会话提交与评测稳定性 (#3346)
* fix(memory): disable unsupported tool and skill extraction

* refactor(memory): retire SessionCompressorV2

* docs: design service import cycle fix

* fix(import): break QueueFS service import cycle

* update

* docs: design memory overview lock coverage fix

* fix(memory): cover overview files in update leases

* docs: design session commit default concurrency 50

* perf(queue): raise session commit concurrency to 50

* docs: revise session commit concurrency design

* docs: plan session commit default 8

* perf(queue): default session commit concurrency to 8

* fix(bot): disable cron during eval chat

* docs: design memory link lock stabilization

* docs: plan memory link lock stabilization

* fix(memory): stabilize link update lock coverage

* docs: cover remapped post-group link locks

* docs: design plain-content patch validation

* docs: design first failing patch diagnostics

* fix: report actual failing patch block

* fix(memory): remap replacement links before locking

* fix(bot): include trusted identity in health probe

* test: consolidate memory contract coverage

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-13 21:48:49 +08:00
Jiahui Zhou 1d02a72b2b Remove qdrant and opengauss vector backends (#3872) 2026-08-07 19:57:45 +08:00
444cc87bf8 feat: OIDC and LDAP as new auth mode for OpenViking (#3708)
* feat: support oidc and ldap auth

* feat: support oidc and ldap auth

* fix: remove heima partner, clean up auth docs, add web-studio unsupported auth banner

- Remove heima from partner list in README (en/zh/ja)
- Remove unsupported env var references (OPENVIKING_AUTH_MODE, OPENVIKING_USERNAME,
  OPENVIKING_PASSWORD) from LDAP auth docs
- Remove temporary switch bash snippets from auth docs
- Fix ldap_password description
- Add web-studio unsupported-auth-mode banner for oidc/ldap servers

* fix: address OIDC/LDAP review comments on auth plugin design

Key changes driven by PR review:

- **Role mapping**: OIDC and LDAP external identities always resolve to
  USER role. Removed map_role() calls and group_membership-based role
  mapping. Admin access is gated by the root API key mechanism only.

- **LDAP credential extraction**: Removed query-parameter-based username/
  password extraction (security concern — passwords in URLs can leak via
  shell history, proxy logs, and monitoring). Clients must use Basic Auth
  header or form data.

- **OIDC identifier sanitization**: Auth0 and other providers may include
  characters like "|" in the `sub` claim. These are now replaced with "_"
  to produce valid OpenViking user identifiers.

- **Dead code removal**: Removed _extract_groups, memberof_attribute,
  require_root_api_key_for_admin, _initialize_api_key_manager, and
  get_request_context_checks from both plugins since they are no longer
  needed.

- **Docs**: Removed query-parameter curl example, memberof_attribute and
  require_root_api_key_for_admin config references.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* feat: support oidc and ldap auth

* feat: support oidc and ldap auth

* fix(auth): bind lazy OIDC imports at module scope

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-06 12:36:37 +08:00
t0saki e910c5feb0 fix(packaging): decouple langchain client from server (#3711) 2026-08-03 21:56:49 +08:00
Hao Zhe b2e1972610 refactor(langchain): extract standalone integration package (#3685)
* refactor(langchain): extract standalone integration package

* fix(langchain): preserve optional legacy imports

* fix(langchain): guard legacy submodule imports
2026-08-03 15:01:13 +08:00
zgy 49b182045b refactor(parser): Refactor code summaries to fixed skeleton-first routing (#3568)
* Refactor code summary skeleton routing

* Simplify code skeleton routing configuration

* Render C tag skeletons as signatures

* Revert "Render C tag skeletons as signatures"

This reverts commit 8e342055f8.

* Simplify fixed code skeleton summary route

* Inline process skeleton rendering

* Simplify code skeleton routing entrypoints

* Fix code summary review issues

* Address final code summary review feedback

* Route failed tags skeletons to LLM fallback

* Restore CUDA and TS extension routing

* Improve code skeleton query coverage

* Route semantic code detection through skeleton support

* Move process skeleton engine into ast package

* Admit skeleton-supported files during directory scan

* Align code summary docs after main merge

* Reduce code skeleton fallback log verbosity

* chore: require grep-ast 0.9.0
2026-07-31 11:38:57 +08:00
t0saki b71f3ed5b3 fix: pin MCP SDK to v1 (#3600) 2026-07-29 14:07:58 +08:00
Rocke Dong ef43b0fc37 fix(deps): restore openviking-sdk lock entry (#3441) 2026-07-22 15:28:24 +08:00
ef4d97ebe3 feat(snapshot): support path-filtered commit history (#3271)
* feat(snapshot): support path-filtered commit history

* refactor(snapshot): reuse SDK git log implementation

* fix(snapshot): harden path-filtered log resource limits

* chore: remove stale SDK lock entry

---------

Co-authored-by: zhanghaoyu.la <zhanghaoyu.la@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-16 19:53:40 +08:00
huangruitengandhuangruiteng 0cf36f483e fix(deps): align bot requests with chardet 7 (#3282)
Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-16 11:09:20 +08:00
Rocke DongandClaude Fable 5 4e142a6b09 chore(deps): sync uv.lock (#3107)
Regenerate uv.lock metadata to match pyproject.toml after two upstream
deps changes that left the lock unreconciled:
- b596e269 (#2965): bump litellm pin <1.89.3 -> <1.90.3
- 47da6ce1 (#3037): merge bot-* extras into a single [bot] extra

Metadata-only diff (13 insertions / 66 deletions) — no resolved package
versions or hashes change. Confirmed by a clean `uv sync` leaving
uv.lock unmodified.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 15:18:25 +08:00
Evo a7a9b71460 fix(deps): bump bot socketio floors for GHSA DoS advisories (#2870) 2026-07-08 18:49:56 +08:00
zgy a50e9fd677 feat: add recursive web crawler based on Scrapy (#2836)
* Refactor recursive web import into HTTP accessor

Move ordinary web page import routing into HTTPAccessor and materialize crawled pages as a temporary directory via WebImporter.

Relocate Scrapy/Playwright crawling under parse.accessors.web_crawler, keep trafilatura extraction inside HTMLParser, and avoid repeated ResourceService.add_resource calls.

Add recursive crawl controls, safe request validation, page/download classification, and focused unit coverage.

* Document recursive web crawler options

* fix(web-crawler): stop SSRF sub-resource block from failing whole render

The playwright fallback validated every sub-resource request against the
SSRF guard and raised on the first disallowed host, failing the entire
page render. volcengine docs load a probe resource on an internal host,
so rendering always failed and the crawler stored the static anti-bot
"Please wait..." challenge page as content.

Now a blocked sub-resource is only aborted; the main document and final
URL still gate the result. Also wait past JS interstitials, retry reads
through in-flight navigation, and reject shell/challenge pages instead of
storing them.

* fix(web-crawler): surface renderer error hint on entry-page failure

When Playwright is unavailable, the renderer returns an actionable install
hint via RenderResult.error, but the spider silently kept the static shell
and WebImporter raised only the generic "Failed to fetch entry page". The
hint never reached the user.

Now the spider records rendered.error on the failed page, and WebImporter
appends the entry page's failure reason to the raised message so the CLI
shows the Playwright install instructions.

* fix(web-crawler): surface render hints and enforce crawl limits

* fix(web-crawler): avoid rendering SSR app pages

* perf(web-crawler): bound render concurrency and cap networkidle wait

Playwright renders were dispatched from parse callbacks without any
concurrency limit, so a page with many child links could spawn dozens of
Chromium pages at once (observed peak 28 for a 20-page crawl), risking OOM
on large sites and starting ~2.3x more renders than needed before
max_pages stopped the crawl. Gate renders with a semaphore sized to
config.concurrency and re-check the success limit after acquiring a slot
so queued callbacks skip rendering once the crawl is already done.

Also cap the networkidle wait at 8s: pages with continuous background
activity (e.g. GraphiQL) never go idle and previously blocked until the
full render timeout, turning a ~3s page into ~38s. Content is ready after
domcontentloaded and _wait_past_challenge covers late-arriving text.

Bump default concurrency 5 -> 10.

* fix(web-crawler): route .html/.htm URLs through recursive WebImporter

An explicit .html/.htm URL is detected as DOWNLOAD_HTML via the extension
map, so access() previously only routed URLType.WEBPAGE to WebImporter and
these URLs fell through to single-file download, silently ignoring
depth/max_pages. Route DOWNLOAD_HTML through WebImporter too, treating a
single-page import as the depth=0 case.

* fix(web-crawler): improve HTML extraction and rendering heuristics

- Drop trafilatura favor_precision=True: it stripped the full body of
  link-dense pages, keeping only headers.
- Only render __NEXT_DATA__ pages with Playwright when their static body
  is too thin; SSR/SSG Next.js pages already ship full text.
- Disable Scrapy telnet console to avoid opening port 6023.

* fix(web-crawler): keep code-hosting single-file URLs off recursive crawler

GitHub/GitLab blob and GitHub raw URLs resolve to a single file, not a
site. Route them through the single-file download path instead of the
recursive WebImporter, which otherwise crawls the hosting UI shell.

* docs(resources): add recursive web crawler usage examples

Add depth/max_pages crawl examples to the HTTP, Python SDK, and CLI
blocks in both the zh and en resource API docs, plus path-prefix
filtering and skip_download_links variants.
2026-07-03 19:25:10 +08:00
yepper 4c1ff1190b Watch sync followup (#2962)
* feat(watch): sync tasks with resource rm and mv

* feat(dependencies): add defusedxml and feedparser packages; update sgmllib3k version

* fix(watch): address review issues in move sync
2026-07-02 20:05:36 +08:00
87329714dd feat(grep): integrate VikingDB bm25 keyword search for grep engine (#2144)
* feat(grep): integrate VikingDB bm25 keyword search for grep engine

* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)

* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison

* fix(schema): upsert data to vikingdb lack of content

* chore: add benchmark for retrieval

* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs

* fix(benchmark): sub uri args; add report

* refactor: code format by ruff

* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf

* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search

* fix: adjust benchmark scripts

* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls

* refactor: new benchmark

* fix: step1 add resource by real code data

* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex

* optimize (benchmark): adjust keywords and ground truth for testing

* fix: truncate 64KB for content field

* optimize: effectiveness add resource plainly

* optimize: change param use of SearchByKeywords from "keywords" to "query"

* optimize(benchmark): refactor effectiveness scripts

* optimize: ensure raw data for content field

* optimize: fulltext analyzer's stop-words only use symbols

* fix: adapt to new ov cli for benchmark

* optimize: reuse file content to avoid re-read AGFS file

* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts

* optimize: benchmark client timeout

* update README

* fix: rm unused param

* fix: default values in docs

* optimize: increase truncate byte size to 1MB for content field for VikingDB

* fix(logger): harden queued stream logging (#2786)

* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock

When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.

During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.

Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.

Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.

Closes: #2752

* fix(logger): harden queued stream logging

---------

Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
2026-06-24 18:46:02 +08:00
Hao Zhe 324f96ebb6 fix(parse): normalize legacy text encodings (#2770)
* fix(parse): normalize text file encodings

* fix(parser): harden text encoding normalization

* fix(parse): normalize text encodings with charset-normalizer

* test(parse): use synthetic gb18030 fixture text

* fix(parse): respect detector rank for non-cjk text

* fix(parse): rescue short simplified chinese text

* fix(parse): preserve korean hanja text

* style(parse): format text encoding tests
2026-06-23 16:55:32 +08:00
Qin Haojie ff258768c2 feat(memory): 引入 User/Peer 记忆隔离模型 (#2236)
* feat(memory): introduce user and peer memory isolation

Unify agent-scoped memory behavior into user-owned memory spaces, add peer_id compatibility for session and retrieval paths, and wire memory_policy through session commit flows.

* feat(memory): align session identity around peer IDs

* feat(search): pass peer id through retrieval

* refactor(memory): remove agent identity from integrations

* fix(memory): isolate peer identity from self extraction

* fix(tau2): provision benchmark user configs

* fix(auth): allow admin keys to access data APIs

* fix(openclaw): enable peer memory policy for peer roles

* fix(openclaw): resolve sender for peer recall

* refactor(session): simplify memory extraction routing

* refactor(ov-cli): reduce formatting-only diff

* refactor(message): remove unused message helpers

* refactor(retrieval): simplify peer target resolution

* refactor(namespace): remove deprecated agent namespace policy

* fix(agent): propagate peer id through integrations

* fix(auth): align integration clients with api-key mode
2026-06-05 10:55:48 +08:00
chenjwandAiden 82225006fb Feat/searchable template (#2193)
* feat: replace searchable memory fields with embedding templates

Use memory-type embedding templates for vectorization, share template rendering with content serialization, and fall back to plain content when embedding rendering cannot be resolved.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* fix(memory): rename role-id isolation config flag

Use role_id_memory_isolation_enabled consistently across config, handler, and tests so the rename matches current prepare_messages behavior.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260522_190730

* auto-commit before eval 20260522_194846

* auto-commit before eval 20260522_200235

* auto-commit before eval 20260522_200720

* update

* auto-commit before eval 20260525_013620
2026-05-25 13:57:26 +08:00
85908e241a Feat/memory link (#2010)
* auto-commit before eval 20260509_181850

* auto-commit before eval 20260509_192618

* update

* auto-commit before eval 20260510_005109

* auto-commit before eval 20260510_011832

* auto-commit before eval 20260510_014114

* auto-commit before eval 20260510_022835

* auto-commit before eval 20260510_025048

* auto-commit before eval 20260510_031034

* auto-commit before eval 20260510_143728

* auto-commit before eval 20260510_172705

* auto-commit before eval 20260510_220133

* auto-commit before eval 20260511_115905

* auto-commit before eval 20260511_121959

* auto-commit before eval 20260511_132120

* auto-commit before eval 20260511_161430

* auto-commit before eval 20260511_163606

* auto-commit before eval 20260511_173943

* auto-commit before eval 20260511_175657

* auto-commit before eval 20260511_224347

* auto-commit before eval 20260511_233109

* auto-commit before eval 20260512_104710

* auto-commit before eval 20260512_111256

* auto-commit before eval 20260512_181905

* auto-commit before eval 20260512_191540

* auto-commit before eval 20260512_192540

* auto-commit before eval 20260512_195710

* auto-commit before eval 20260513_000746

* auto-commit before eval 20260513_004221

* auto-commit before eval 20260513_004656

* refactor: migrate logger calls to tracer in extract_loop modules

Replace logger.warning/error/info with tracer.error/info in extract_loop
related modules for better observability (console + OpenTelemetry spans).

Modules updated:
- agent_experience_context_provider.py (5 replacements)
- extract_loop.py (4 replacements)
- memory_updater.py (9 replacements)
- session_extract_context_provider.py (4 replacements)
- utils/json_parser.py (7 replacements)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* auto-commit before eval 20260513_123007

* auto-commit before eval 20260513_125305

* auto-commit before eval 20260513_135421

* auto-commit before eval 20260513_141013

* auto-commit before eval 20260513_143455

* auto-commit before eval 20260513_145401

* auto-commit before eval 20260513_163345

* auto-commit before eval 20260514_105906

* auto-commit before eval 20260514_112912

* auto-commit before eval 20260514_120308

* auto-commit before eval 20260514_122022

* auto-commit before eval 20260514_134800

* auto-commit before eval 20260514_135615

* auto-commit before eval 20260514_135818

* auto-commit before eval 20260514_142941

* auto-commit before eval 20260514_162401

* auto-commit before eval 20260514_231859

* auto-commit before eval 20260515_104122

* auto-commit before eval 20260515_122140

* auto-commit before eval 20260515_122942

* auto-commit before eval 20260515_144941

* auto-commit before eval 20260515_154736

* auto-commit before eval 20260515_181643

* auto-commit before eval 20260515_182727

* auto-commit before eval 20260515_183056

* auto-commit before eval 20260515_183652

* auto-commit before eval 20260515_183825

* auto-commit before eval 20260515_202731

* auto-commit before eval 20260516_001144

* auto-commit before eval 20260516_011749

* auto-commit before eval 20260516_015903

* auto-commit before eval 20260516_020505

* auto-commit before eval 20260516_130701

* auto-commit before eval 20260516_144342

* auto-commit before eval 20260516_151043

* Harden memory graph rendering and patch guidance.

Escape embedded graph data for script safety, add a vis-network load guard, tighten graph layout defaults, and clarify SEARCH guidance so patch content stays bound to the target file/page context.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260517_005258

* auto-commit before eval 20260517_012903

* auto-commit before eval 20260517_014036

* auto-commit before eval 20260517_015726

* auto-commit before eval 20260517_024952

* auto-commit before eval 20260517_032518

* auto-commit before eval 20260517_135114

* auto-commit before eval 20260517_143238

* auto-commit before eval 20260517_154858

* auto-commit before eval 20260517_200556

* auto-commit before eval 20260517_215025

* fix: keep memory storage plain and render graph links on display

Store memory bodies as plain text in VikingFS and move link rendering to graph display so repeated writes no longer persist nested markdown links. Also tighten link renderer path handling so cross-user relative paths are rejected and strip_links preserves viking and absolute targets.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260518_001945

* fix: invert selected graph node colors

Make the currently selected memory node use a light background with dark text so it stands out against the dark graph theme.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260518_011327

* update

* auto-commit before eval 20260518_161813

* auto-commit before eval 20260518_165104

* auto-commit before eval 20260518_174259

* update

* auto-commit before eval 20260518_224834

* auto-commit before eval 20260518_233319

* auto-commit before eval 20260518_235712

* auto-commit before eval 20260519_135952

* fix memory patch failure logging

Keep dry-run patch validation from emitting a misleading patch_handler warning, and record skipped field updates from MemoryUpdater where the failure is handled.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260519_213142

* fix(memory): fan out links for shared page ids

Expand _resolve_links so shared page ids resolve across every operation URI instead of collapsing to a single path. Align the page-id and extract-loop tests with the current API contract and the multi-URI link behavior.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260520_195141

* auto-commit before eval 20260520_215911

* auto-commit before eval 20260520_222335

* update

* style(memory): clean up formatter drift

Apply the remaining formatter-driven cleanup in the memory modules so the working tree stays clean before the next behavior changes. This keeps helper signatures and string literals aligned with current lint output.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260521_130517

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-21 14:58:24 +08:00
yepper ddcd3fb9c8 chore(format): align python and c++ file formatting (#2001)
* chore(format): align python and c++ file formatting

* chore: update urllib3 to 2.7.0 and clean test imports

1. bump urllib3 dependency from 2.6.3 to 2.7.0
2. remove unused pytest import and RoleScope import from test file

* style: format list comprehensions and lambda function for readability

Adjust the line breaks in the list comprehension in the VikingSearchTool class to follow standard Python formatting conventions, and rewrap the lambda assignment in the test case to improve code readability without changing functionality.

* style: fix line wrapping and remove extra blank line

- remove stray blank line in ov_server.py
- wrap long logger.info line in memory.py for better readability

* style: fix targeted ruff lint violations

* chore: clean up unused imports and reorder code

This commit removes unused imports, reorders import statements for better consistency,
and simplifies some test file imports. Changes include:
- Remove redundant blank lines and unused imports across multiple test files and core modules
- Reorder imports in openviking hooks module to follow standard layout
- Fix import ordering in memory isolation handler
- Simplify php parser type imports
- Move volcengine mock import to correct position in test file

* refactor(uri utils): remove extra blank lines in uri.py

clean up redundant whitespace to improve code readability
2026-05-13 17:53:09 +08:00
Hao Zhe 81d1b5afd7 feat(langchain): add LangChain and LangGraph context adapters (#1964)
* feat(langchain-langgraph): add adapter primitives

* feat(langchain-langgraph): add context backend lifecycle

* fix(langchain-langgraph): harden context backend integration

* fix(langchain-langgraph): accept canonical store result URIs

* docs(langchain): point users to runnable examples

* docs(langchain): add missing integration examples

* refactor(langchain): address integration review feedback

* fix(langchain): address review-blocking integration bugs

* fix(langchain): honor user ids and safe store filters

* fix(langchain): reject unsupported store TTL writes
2026-05-12 16:28:11 +08:00
99ae8fea71 Feat/memory isolation (#1905)
* feat(session): add account namespace policy and shared sessions

Unify namespace resolution across filesystem, indexing, and session storage.
Add account-shared session paths, role_id auth semantics, and an HTTP demo
script for the four namespace-policy combinations.

* update

* update

* fix: add missing parse_iso_datetime import

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix: add registry parameter to apply_operations method

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* update

* fix: memory extraction bug fixes

- session_service: add archive_uri for manual_extract
- memory_type_registry: raise error on duplicate memory_type
- schema_model_generator: use List for all memory_types
- memory_updater: fix current_value issue, use List[List[Message]] for MessageRange
- memory_isolation_handler: iterate nested message groups
- compressor_v2: strip MEMORY_FIELDS comment in memory_diff.json

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* update

* fix: remove archive_uri from _generate_archive_summary_async call

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix: allow custom templates to override built-in memory types

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore: increase default timeout to 10 minutes in import_to_ov

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore: increase embedding slow call warning threshold to 3s

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat: add format error retry in extract loop (max 1 retry)

- Add format error message to guide LLM output correct JSON
- Allow max 1 retry per extraction run
- Increment max_iterations to accommodate retry

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix: poll for task completion in test_group_chat.py

- Wait for memory extraction task to complete before verification
- Wait for vectorization to complete via wait_processed

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat: memory isolation for multi-user group chat and single-chat mode

Add multi-user memory isolation support: search_memory now accepts
multiple user_ids and merges results across user spaces. Memory
isolation handler prioritizes role_ids from messages over ctx defaults,
and schema_model_generator skips user_id/agent_id fields for
message-based memory types (events with ranges).

Add --single-chat and --memory-user CLI flags throughout the stack
(commands, channels, context, agent loop) to control whether
role_id/speaker is set on messages and which users' memories are
retrieved. Benchmark scripts (import_to_ov, run_eval,
import_and_eval_one) support both modes.

Revise memory templates: profile adds multi-person schema with
emotional needs and polarity; preferences adds category grouping and
inferred-preference tagging; events enforces one-event-per-file with
time_references and causal_relationships fields; identity/soul remove
redundant merge_op fields. Also add extraction_enabled config option
and fix minor logging issues.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* update

* refine events extraction prompts and remove start_line from SearchReplaceBlock

- Update events.yaml: simplify extraction rules, emphasize comprehensiveness,
  indirect speech conversion, narrative style, and avoid splitting single events
- Remove start_line field from SearchReplaceBlock (unused by LLM)
- Use getattr with fallback in patch_handler for backward compatibility
- Update tests to match new SearchReplaceBlock interface

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* auto-commit before eval 20260426_005734

* auto-commit before eval 20260426_010358

* auto-commit before eval 20260426_014944

* auto-commit before eval 20260426_022120

* auto-commit before eval 20260426_161616

* auto-commit before eval 20260426_184659

* auto-commit before eval 20260427_001552

* auto-commit before eval 20260427_121245

* auto-commit before eval 20260427_144044

* auto-commit before eval 20260427_153412

* auto-commit before eval 20260427_221932

* update

* auto-commit before eval 20260427_235724

* auto-commit before eval 20260428_122102

* auto-commit before eval 20260428_135117

* auto-commit before eval 20260428_152454

* Remove obsolete ov_dream skill files

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* auto-commit before eval 20260429_003429

* auto-commit before eval 20260429_010413

* docs: add memory link design spec

Move and update memory link design document: Dream定时整理 → 外部Bot T+1触发整理.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs: add memory_diff.json structure to session.commit docs

Document the memory_diff.json file written during session.commit Phase 2,
including JSON schema, field descriptions, and storage structure updates
in both concept and API docs (en/zh).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* auto-commit before eval 20260429_150714

* auto-commit before eval 20260429_210907

* auto-commit before eval 20260429_211246

* auto-commit before eval 20260429_215224

* auto-commit before eval 20260429_215608

* auto-commit before eval 20260429_220349

* auto-commit before eval 20260429_223650

* auto-commit before eval 20260430_153938

* auto-commit before eval 20260430_213747

* auto-commit before eval 20260430_225924

* auto-commit before eval 20260501_010556

* auto-commit before eval 20260501_032425

* auto-commit before eval 20260501_142823

* auto-commit before eval 20260501_163459

* auto-commit before eval 20260501_172516

* auto-commit before eval 20260501_173356

* auto-commit before eval 20260501_174039

* auto-commit before eval 20260501_180151

* auto-commit before eval 20260501_183653

* auto-commit before eval 20260501_185122

* auto-commit before eval 20260501_190512

* auto-commit before eval 20260501_213556

* auto-commit before eval 20260501_222537

* auto-commit before eval 20260501_224527

* auto-commit before eval 20260501_233947

* auto-commit before eval 20260502_013834

* auto-commit before eval 20260502_015557

* auto-commit before eval 20260502_021439

* auto-commit before eval 20260502_023732

* auto-commit before eval 20260502_155020

* auto-commit before eval 20260503_005125

* auto-commit before eval 20260503_005457

* auto-commit before eval 20260503_010359

* auto-commit before eval 20260503_011801

* auto-commit before eval 20260503_013354

* auto-commit before eval 20260503_015552

* auto-commit before eval 20260503_234630

* auto-commit before eval 20260503_235842

* auto-commit before eval 20260504_000706

* auto-commit before eval 20260504_001553

* update

* auto-commit before eval 20260504_081658

* auto-commit before eval 20260504_082343

* auto-commit before eval 20260506_111759

* auto-commit before eval 20260506_165607

* auto-commit before eval 20260506_214411

* auto-commit before eval 20260507_151946

* auto-commit before eval 20260507_153237

* auto-commit before eval 20260508_110748

* auto-commit before eval 20260508_111642

* fix

* fix

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-08 14:52:27 +08:00
sentisso 9b45c21499 feat(upload): respect root and nested .gitignore during filtering (#1812)
* feat: adding .gitignore compliance

* docs

* fix

* revert
2026-04-30 14:01:54 +08:00
t0saki ccad9c5e0e feat(server): native MCP endpoint with 9 tools aligned to VikingBot (#1738)
* feat(server): add native MCP endpoint at /mcp

Serve 5 MCP tools (search, read, store, forget, health) directly from
the OV FastAPI server via streamable HTTP transport. This eliminates the
need for the Node.js MCP subprocess — the plugin's .mcp.json now points
to the server URL instead of spawning a process.

Identity headers (X-OpenViking-Account/User/Agent) are propagated to
service-layer calls via contextvars ASGI middleware.

* fix(mcp): disable DNS rebinding protection for reverse proxy compatibility

MCP SDK auto-enables host validation for localhost, rejecting requests
with external Host headers (e.g. from Cloudflare/Nginx reverse proxy).

* fix(mcp): reuse auth.resolve_identity for MCP endpoint authentication

MCP endpoint previously had no authentication — requests fell through
with default/default identity. Now delegates to the same resolve_identity
used by all REST routes, so auth_mode, API key validation, and identity
resolution are handled identically.

* fix(mcp): fix import path for TextPart in store tool

openviking.session.parts does not exist; the correct module is
openviking.message.part.

* fix(mcp): store tool now creates a new session and commits immediately

Each store call creates a unique session, adds the message, and commits
right away so memories are extracted and searchable without waiting for
a token threshold.

* chore: add mcp>=1.27.0 dependency for native MCP endpoint

* fix(mcp): align search/forget tools with REST API, fix forget crash

- Remove SEARCH_TARGETS and per-scope loop; use single
  service.search.find(target_uri="") call matching REST API behavior
- Fix forget crash: FSService has no delete(), use rm() instead
- Replace fragile _is_memory_uri() substring check with ContextType
- search tool: replace scope param with target_uri for direct passthrough
- Work directly with FindResult/MatchedContext objects instead of
  dict-munging via to_dict()

* fix(mcp): fail-closed on missing identity, remove unused Role import

- _get_ctx() now raises UnauthenticatedError instead of defaulting to
  ROOT when identity contextvar is not set
- Remove unused Role import
- Clean up comments in create_mcp_app

* test(mcp): add unit tests for MCP endpoint tools

17 tests covering all 5 MCP tools and identity propagation:
- _get_ctx: returns context when set, raises UnauthenticatedError when not
- health: healthy/unhealthy responses
- search: no results, with resource, with target_uri
- read: nonexistent URI, directory listing, batch reads
- store: user and assistant roles
- forget: input validation, non-memory guard, URI deletion, query fallback
- Route registration: /mcp route exists in app

* docs(mcp): update integration guide with verified platforms and correct tools

- Add verified platforms table (Claude Code, ChatGPT/Codex, Claude.ai,
  Manus, Trae)
- Document authentication (X-Api-Key / Bearer token)
- Add Claude.ai OAuth proxy (MCP-Key2OAuth) instructions
- Update tool table to match actual implementation (search, read, store,
  forget, health) — remove stale tool names
- Reorganize client config: generic first, then platform-specific

* feat(mcp): expand to 7 tools aligned with vikingbot, split read/list

Align MCP tool surface with vikingbot/agent/tools/ov_file.py:

- Split read/list: read is file-only with semaphore(10) concurrency;
  list is directory-only with recursive support
- store: accept batch messages[] (was single text), matching
  VikingMemoryCommitTool
- search: add min_score parameter (default 0.35), matching
  VikingSearchTool
- add_resource: new tool for adding files/URLs to resources
- Use @mcp.tool(name="list") to avoid shadowing Python builtin

7 tools: search, read, list, store, add_resource, forget, health

* feat(mcp): add grep and glob tools, update docs to 9 tools

Add grep (multi-pattern regex search) and glob (file pattern matching)
MCP tools to align with VikingBot's full tool surface. Update EN/ZH
integration docs to reflect all 9 tools with correct parameters.

* fix(mcp): store schema, forget safety, remove memories-only restriction

- store: use Pydantic StoreMessage model so MCP schema includes
  required role/content field definitions (was bare dict[str, str])
- forget: remove query parameter entirely — deletion requires exact URI,
  use search tool first to find candidates
- forget: remove /memories/ path restriction, allow deleting any URI

* docs(mcp): update forget tool description — exact URI only, no query

* fix(mcp): rename list_dir to ls, add forget safeguard, use Bearer in docs

- Rename list_dir → ls (MCP tool name stays "list") to avoid confusion
  with "only lists directories"
- Add safeguard to forget tool description: irreversible, requires user
  confirmation
- Docs: use Authorization: Bearer in all examples (standard, consistent
  with OAuth proxy flow)
- Fix ruff format on mcp_endpoint.py and test_mcp_endpoint.py
2026-04-28 15:19:16 +08:00
chenjw 9bcbaad0ed Feat/memory diff (#1710) 2026-04-25 21:19:10 +08:00
baojun-zhangandMaojiaSheng 17d2c5603e feat(observability): unify observability context && support otel && etc. (#1666)
* feat(observability): unify OTLP metrics export, log/trace context, and telemetry bridging
- - Add OTLP metrics http/grpc exporter that pushes MetricRegistry snapshots
- - Decouple telemetry response payload from telemetry collection; always finish() and bridge summary to metrics
- - Unify observability config under server.observability (metrics/traces/logs siblings); update ov.conf.example and docs (zh/en)
- - Improve log/trace correlation via structured context injection
- - Add/adjust tests for exporter lifecycle, config loader, metrics/telemetry runtime
- BREAKING CHANGE: remove legacy telemetry.* config path; use server.observability.*

* feat(observability): import Status/StatusCode for LogToSpanEventFilter

* feat(observability): fix check issue

* feat(observability): format code

---------

Co-authored-by: MaojiaSheng <shengmaojia@bytedance.com>
2026-04-24 21:32:25 +08:00
Hao ZheandZayn Jarvis 01403312ea feat(vlm): add Codex, Kimi, and GLM VLM support (#1444)
* feat(vlm): add Codex OAuth-backed VLM setup and docs

* fix(codex): address PR review follow-up issues

* feat(vlm): add Kimi and GLM backends

* refactor(vlm): simplify codex auth flow and docs

* fix: update code comments and doctor validation

* chore: update uv.lock after merge

* Refine Codex auth flow and VLM backend integrations

* Take over mirrored Codex auth on refresh

* feat(vlm): refine provider setup and auth flow

* style: format VLM and setup files

* style: fix lint import ordering

* fix(codex): harden auth refresh and disable streaming

* fix(codex): translate tool history and refresh auth safely

* style(lint): fix changed-file ruff violations

* fix(init): refine cloud VLM setup prompts

* style(lint): format setup wizard changes

---------

Co-authored-by: Zayn Jarvis <zhiheng.liu@bytedance.com>
2026-04-22 11:01:23 +08:00
chenjwandClaude Opus 4.6 c430b97493 Fix/openclaw addmsg (#1391)
* fix(ragfs): localfs write should respect WriteFlag to truncate file

The localfs write implementation was ignoring the WriteFlag parameter and
not truncating the file when WriteFlag::Create or WriteFlag::Truncate was
specified. This caused files to be appended instead of overwritten when
writing empty content.

Root cause: The _flags parameter was declared but never used. When opening
an existing file, .truncate(true) was not called.

Fix: Use OpenOptions::new().truncate(should_truncate) based on the flags.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(session): clear messages.jsonl after commit and improve memory extraction

- Add workaround to delete messages.jsonl before writing to work around
  RAGFS localfs not truncating files (now fixed in ragfs)
- Add _get_latest_archive_last_msg_time to filter messages for memory
  extraction by timestamp
- Fix memory extraction to only process new messages after last archive

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* update

* fix(openclaw): add HEARTBEAT filtering in sanitizeUserTextForCapture

过滤 HEARTBEAT 健康检查消息,避免 HEARTBEAT.md 和 HEARTBEAT_OK
进入 session 存档。

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* revert: restore logger.warn instead of warnOrInfo

Reverting to use logger.warn directly. The warnOrInfo wrapper was
unnecessary as logger always has warn method.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore(session): remove RAGFS truncate workaround

The ragfs localfs write now respects WriteFlag to truncate files,
so the workaround to delete messages.jsonl before writing is no
longer needed.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(openclaw): unify addSessionMessage with parts API

- Remove old addSessionMessage (string content)
- Keep addSessionMessage but accept parts array
- Update index.ts and context-engine.ts to use new format
- Update tests to check parts structure

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(openclaw): use optional chaining for logger.warn

TypeScript strict mode requires handling possibly undefined warn method.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(openclaw): log trace_id from commit results

Add trace_id to CommitSessionResult type and log it for all
commitSession calls (afterTurn, commitOVSession, compact).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(openclaw): add tool_input to tool parts

- Add toolInput field to ExtractedMessage type
- Extract tool_input from msg.toolInput in extractNewTurnMessages
- Pass tool_input when constructing tool parts for addSessionMessage

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(openclaw): lookup toolInput from preceding toolUse messages

When extracting toolResult messages, look up the corresponding toolInput
from the preceding assistant messages' toolUse blocks using toolCallId/
toolUseId as the key.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(openclaw): add more field name variations for toolInput/toolCallId

Support additional field names: arguments, toolInput, tool_call_id

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(openclaw): lookup toolInput from preceding toolCall messages

- Scan all messages to find toolCall/toolUse/tool_call blocks
- Extract tool input from arguments/input/toolInput fields
- Match toolResult with toolCall using toolCallId/toolUseId

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-14 14:03:04 +08:00
Jiahui Zhou b174deb410 Fix ci (#1307)
* fix(ci): fix build ci

* build: move ragfs-python packaging into setup.py
2026-04-09 00:05:18 +08:00
Jiahui Zhou 0751a11ae4 fix(lark): add lark-oapi (#1285) 2026-04-08 11:06:20 +08:00
chenjw 7f05828f53 Feature/memory opt (#1159) 2026-04-06 15:50:18 +08:00
MaojiaShengandopenviking dc052ce3e3 reorg: Rewrite agfs to ragfs with rust (#1221)
* reorg: rewrite agfs with rust, and named with ragfs, keep License

* reorg: rewrite agfs with rust, and named with ragfs, keep License

* reorg: rewrite agfs with rust, and named with ragfs, keep License

* reorg: rewrite agfs with rust, and named with ragfs, keep License

* reorg: rewrite agfs with rust, and named with ragfs, keep License

* reorg: rewrite agfs with rust, and named with ragfs, keep License

* reorg: rewrite agfs with rust, and named with ragfs, keep License

* reorg: rewrite agfs with rust, and named with ragfs, keep License

* fix: grep level limit

* fix: grep root

* fix: import error

* fix: rust code optimazation

* fix: CI error

* fix: CI go mod cache

* fix: grep level limit

* fix: CI

---------

Co-authored-by: openviking <openviking@example.com>
2026-04-05 21:17:45 +08:00
Jiahui Zhou 95c29de4bf fix(ci): refresh uv.lock for docker release build (#1094) 2026-03-30 20:44:01 +08:00
Qin Haojie 51a1e60cae fix(security): restore litellm integrations below 1.82.6 (#966)
Re-enable LiteLLM-backed providers and tools after the temporary hard-disable,
while pinning the dependency below 1.82.6 to avoid the compromised release window.
2026-03-25 18:11:18 +08:00
Qin HaojieandClaude Opus 4.6 9eb6a5a2b8 fix(security): hard-disable litellm integrations (#937)
Disable LiteLLM-backed providers, config entry points, and image tooling so
new installs and runtime config cannot route through a compromised dependency.

Co-Authored-By: Claude Opus 4.6
2026-03-24 23:12:40 +08:00
chenjwandClaude Opus 4.6 2771765298 Refactor memory extract (#916)
* docs: add memory extractor templating and update mechanism optimization design document

- Add bilingual (English/Chinese) design document for memory templating system
- Include YAML-based MemoryTypeRegistry with 8 built-in types
- Detail ReAct 3+1 phase flow with pre-fetch optimization
- Describe 3-operation Schema: write/edit/delete
- Document RoocodePatch SEARCH/REPLACE format
- Explain dual-mode design: simple mode vs template mode
- Cover pre-fetch optimization: ls directories + read .abstract.md/.overview.md + search once
- Include merge operations: patch, sum, avg, immutable
- Address #578: allow custom prompt template addition and specification

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add memory templating system with ReAct orchestrator

- Add YAML-configurable memory schemas (cards, events, entities, etc.)
- Implement MemoryReAct with tool use (read/find/ls)
- Add schema-driven memory operations (write_uris/edit_uris/delete_uris)
- Implement memory patch handler for incremental updates
- Add comprehensive test suite

* refactor: memory extractor templating system with ReAct orchestrator

## Summary

Implement memory templating system (GitHub Issue #578) - a complete
rewrite of the memory extractor subsystem to support YAML-configurable
memory types instead of hardcoded categories.

## Key Changes

### Architecture
- Replace hardcoded 8 memory types with YAML-configurable schema system
- Add MemoryTypeRegistry to load memory type definitions from YAML files
- Dynamic Pydantic model generation from schema for type safety
- Field-level merge operations: PATCH, SUM, IMMUTABLE

### Memory Extraction Flow
- Implement ReAct orchestrator for single-pass memory updates
- MemoryUpdater for applying operations to storage
- Memory tools (read, search, ls) for ReAct loop
- Stable JSON parser with 5-layer fault tolerance

### File Naming & Storage
- Semantic filenames from template ({topic}.md instead of random IDs)
- Two memory modes: simple mode and template mode
- MEMORY_FIELDS HTML comment for structured metadata

### Configuration
- 9 YAML templates in openviking/prompts/templates/memory/
- memory_config.py for memory system configuration
- Dual-threshold compact upload mechanism in design doc

### Deletions
- Remove old memory_content.py, memory_data.py, memory_operations.py
- Remove memory_types.py, memory_utils.py, memory_patch.py
- Remove corresponding old test files

### Updated Components
- VLM backends (litellm, openai, volcengine) for new interfaces
- Session and service core integration
- Test suite for new architecture

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: pass ctx/user/session_id in commit_async for memory extraction

## Summary

Fix missing parameters in commit_async() when calling extract_long_term_memories().
The synchronous commit() method correctly passes these parameters, but the async
version was missing them, causing memory extraction to be skipped.

## Changes

- Pass user=self.user, session_id=self.session_id, ctx=self.ctx
  in commit_async() when calling extract_long_term_memories()

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: convert FindResult to dict before returning from search tool

## Summary

Fix JSON serialization error by converting FindResult object to dict
using its to_dict() method before returning from MemorySearchTool.

## Changes

- In MemorySearchTool.execute(), return search_result.to_dict()
  instead of search_result directly

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: swap None check before accessing final_operations in memory_react

Also rename schema_models.py to schema_model_generator.py for clarity.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add edit_overview support and optimize memory registry initialization

- Add edit_overview_operations to MemoryUpdater for updating .overview.md files
- Optimize MemoryTypeRegistry initialization in SessionCompressorV2 (load once)
- Various memory templating system improvements

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: remove unnecessary indent in JSON schema output

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: add markdown link format hint to overview field description

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add pre-fetch search based on user messages in conversation

Also fix duplicate line in system prompt.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* rebase

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-24 15:28:46 +08:00
Jiahui Zhou 033826f028 Migrate vectordb engine to abi3 packaging (#897)
chore: remove pybind11 remnants after abi3 migration

fix: scope linux abi3 wheel smoke test to engine loader
2026-03-23 19:11:07 +08:00
baojun-zhang 8cdecd9163 feat: Add multi-tenant file encryption capability (#828)
* feat: Add multi-tenant file encryption capability

* feat: move cli command `ov crypto` to `ov system crypto` && ruff format && add encryption config example

* feat: reformat code

* feat: adjust crypto.rs to support diff platform

* feat: reformat code
2026-03-21 13:35:46 +08:00
chethanuk 33f1a8f42b feat(gemini): add GeminiDenseEmbedder text embedding provider (#751)
* feat(gemini): add GeminiDenseEmbedder with is_query routing per #702 pattern

- GeminiDenseEmbedder: accept query_param/document_param, use is_query
  in embed() and embed_batch() to select task_type at call time
- EmbeddingConfig: add Gemini provider, factory, validation, dimension
- No get_query_embedder/get_document_embedder/_get_contextual_embedder
  (removed in #702; embed(is_query=True/False) is the pattern)
- Tests use embed(text, is_query=True/False) pattern throughout
- Rebased onto current upstream/main

* fix(gemini): remove task_type config field, fix conditional import for CI

- Remove task_type from EmbeddingModelConfig (query_param/document_param suffice)
- Wrap GeminiDenseEmbedder import in try/except (google-genai is optional)
- Update tests for removed field
2026-03-20 21:26:23 +08:00
Qin Haojie ae35f46c75 Revert "feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF) (#607)" (#703)
This reverts commit 95bd19797f.
2026-03-17 21:29:21 +08:00
chethanuk 95bd19797f feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF) (#607)
* feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF)

Native text + multimodal (image, video, audio, PDF) embedding via `gemini-embedding-2-preview` (google-genai 1.67.0). Additive provider pattern — Volcengine remains the default; Gemini is opt-in via `provider: "gemini"` in `ov.conf`.
    - **Model**: `gemini-embedding-2-preview`
    - **Input**: text, image, video, audio, PDF (17 MIME types)
    - **Output dimension**: 128–3072 (default: **3072**, recommended: 768 / 1536 / 3072)
    - **Input token limit**: **8,192 tokens**
    - **Supported MIME types**: `image/jpeg`, `image/png`, `image/gif`, `image/webp`, `audio/mpeg`, `audio/mp3`, `audio/wav`, `audio/ogg`, `audio/flac`, `video/mp4`, `video/mpeg`, `video/mov`, `video/avi`, `video/webm`, `video/wmv`, `video/3gpp`, `application/pdf`
- Gemini Embedding 2 Multimodal Support: Introduced a new GeminiDenseEmbedder to support native text and multimodal (image, video, audio, PDF) embedding using the gemini-embedding-2-preview model. This is an opt-in provider via configuration.
- Extended Queue Pipeline for Multimodal Content: The EmbeddingMsg now carries media_uri and media_mime_type to facilitate multimodal content processing. The TextEmbeddingHandler.on_dequeue() method was updated to read raw bytes from viking_fs and call embed_multimodal() when applicable.
- End-to-End Configuration and Security: The EmbeddingConfig now registers the 'gemini' provider with a task_type field. A critical security validation was added to ensure media_uri matches context_data['uri'] before file reads, preventing forged queue messages from accessing arbitrary files. If validation fails or multimodal embedding fails, it falls back to text embedding.
- Multimodal Content Representation: A new ModalContent dataclass was introduced to represent media references, including MIME type, URI, and optional raw data, enabling the Vectorize object to encapsulate both text and media for embedding

* feat: Add asynchronous batch embedding with concurrency control to Gemini embedder.

* Reduce scope to use GeminiDenseEmbedder as only text embed
2026-03-17 21:15:12 +08:00
Lam Ngoc Nguyen 559eef38c9 feat(parse): add support for legacy .doc and .xls file formats (#652)
* feat(parse): add support for legacy .doc and .xls file formats

Add LegacyDocParser using olefile to extract text from Word 97-2003
binary .doc files via OLE2 stream parsing with piece table support
and multi-level fallbacks.

Extend ExcelParser to handle .xls files using xlrd, branching the
parse logic based on file extension while keeping openpyxl for
.xlsx/.xlsm.

New dependencies: olefile>=0.47, xlrd>=2.0.1

* fix(parse): handle date and boolean cell types in xlrd .xls parsing

Check cell.ctype for XL_CELL_DATE and XL_CELL_BOOLEAN to avoid
outputting raw float serial numbers for dates and numeric 0/1 for
booleans.

* fix(parse): harden legacy .doc/.xls parsers per code review

excel.py:
- Enable formatting_info=True so xlrd detects date cells properly
- Add on_demand=True and release_resources() for memory efficiency
- Handle all xlrd cell types: DATE (with time), BOOLEAN, ERROR, BLANK, EMPTY
- Display integers without trailing .0
- Extract cell formatting to _format_xls_cell static method

legacy_doc.py:
- Add 50MB stream size cap to prevent DoS from crafted files
- Cap ccpText at 10M chars to prevent memory exhaustion
- Add FIB version check (require Word 97+ / nFib >= 0x00C1)
- Add minimum buffer length check before struct.unpack_from
- Fix Grpprl skip loop to prevent spin on zero-length entries
- Add _clean_word_text for \x0B (soft break) and \x0C (section break)
- Log warnings for pieces extending beyond stream bounds
- Cap fallback extract to max stream size
2026-03-16 16:29:51 +08:00
MaojiaShengandopenviking 0fb8f13a13 fix: windows zip path, code repo indexing, search retrieval, account id, rust cli version... (#577)
* fix: windows zip path norm

* fix: account id in vector db

* fix: add some log

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

* fix: add some log, and fixed search

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-15 10:04:51 +08:00
chuanbao666 84d206dd7a fix(agfs): refactor Makefile for cross-platform compatibility (#571)
* fix(agfs): refactor Makefile for cross-platform compatibility

* build: support python 3.14
2026-03-14 17:45:11 +08:00
Qin HaojieandClaude Opus 4.6 1b31130389 fix(openclaw-memory-plugin): add missing files to download list and merge config properly (#516)
Add tsconfig.json and new source files (client.ts, process-manager.ts,
memory-ranking.ts, text-utils.ts) to install scripts. Merge plugin load
paths and allow list instead of overwriting existing config. Adjust
default recallLimit to 6. Update uv.lock with resolved bot-full deps.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 20:09:55 +08:00
chenjwandDuTao c4cdb5fb1b Fix/bot Fix Memory Consolidation Issues and Integrate Vikingbot as OpenViking Optional Dependency (#492)
* fix

* add log

* add log

* update

* Opt memory

* Opt memory

* Opt memory

* Opt memory

* Fix memory consolidation issues: duplicate session write, hook error, and async

- Fix _consolidate_memory hook error by creating temporary Session when given message list
- Remove duplicate session save logic from _consolidate_memory (handled by caller)
- Make normal memory consolidation async like /new command
- Other fixes: streaming response in CLI and OpenAPI channel

* Integrate vikingbot as openviking optional dependency

- Add vikingbot dependencies to pyproject.toml optional-dependencies
- Update install prompts from vikingbot[X] to openviking[bot-X]
- Update README installation instructions
- Configure setuptools to find vikingbot in bot/ directory
- Add vikingbot package data and script entry

* Remove bot/pyproject.toml

vikingbot is now integrated as part of openviking, no need for separate pyproject.toml

* Update vikingbot installation instructions in root READMEs

- Update from 'uv pip install -e bot/' to 'uv pip install -e ".[bot]"'
- Add quotes around openviking[bot] for shell safety

* update

* Use singleton pattern for VikingClient in hooks

Cache VikingClient instances by workspace_id to avoid repeated client creation
- Add _client_cache dictionary for caching
- Add get_cached_client() helper function
- Update both hooks to use cached client

* Use global singleton for VikingClient instead of per-workspace cache

- Change from per-workspace_id cache to true global singleton
- Create client with None as agent_id (works for all workspaces)
- Simplify client management

* uv run ruff format

---------

Co-authored-by: DuTao <dutao.1786@bytedance.com>
2026-03-09 22:11:03 +08:00
MaojiaShengandopenviking cb30ab7892 fix: add-resource --to and --parent (#475)
* fix: github zip download timeout

* feat: add-resource --to and --parent args modified

* fix: grep for binding-client

* fix: --to --parent

* fix: --to --parent

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-09 17:58:31 +08:00
MaojiaShengandopenviking c4209da6cc chore: downgrade golang version limit, update vlm version to seed 2.0 (#425)
* feat: define a system path for future deployment

* feat: define a system path for future deployment

* fix: golang downgrade to 1.19

* fix: golang downgrade to 1.19, and change doc

* docs: change model recommendation

* fix: golang downgrade to 1.19, and change doc

* fix: mv test files

* fix: loguru

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-05 14:34:02 +08:00
MaojiaShengandopenviking e4b1ffed10 Feat: CLI optimization (#389)
* feat: remove python cli and disable python -m openviking

* feat: ls, tree, find, search, grep all use --node-limit as the result limiting arg

* feat: add ls -n

* feat: update agfs to support grep -n

* feat: update agfs to support grep -n

* docs: cancel modify

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-03 13:23:37 +08:00
chuanbao666 ac17311b55 fix(agfs): agfs sdk默认安装从本地安装 (#355) 2026-02-28 17:58:05 +08:00