Add WebFeedAccessor (priority 60) that turns a single sitemap /
sitemapindex / RSS / Atom URL into ONE resource tree: it mirrors every
listed page into a temp directory and reuses the existing DirectoryParser
pipeline (the same "fetch-many -> dir -> tree" contract as GitAccessor).
A watch on the feed URL keeps the whole site refreshed (new pages added,
removed pages dropped on each rebuild).
- New openviking/parse/accessors/web_feed_accessor.py: WebFeedAccessor +
sitemap/feed extractors (nested sitemapindex recursion with depth cap,
RSS 2.0 / Atom via feedparser), bounded concurrent polite mirroring,
robots.txt, same-host / include / exclude / max_pages limits.
- args={"site": true} forces whole-site ingestion from a bare domain or
page by auto-discovering the sitemap/RSS (robots.txt, HTML
<link rel=alternate>, conventional paths); {"site": false} opts a
feed-looking URL back out to HTTPAccessor.
- Thread accessor-selection kwargs through can_handle; the registry
tolerates accessors whose can_handle lacks **kwargs (back-compatible).
- Single-page adds get a non-blocking "this site exposes a sitemap/RSS"
suggestion appended to the MCP add_resource response, gated to the
site root only; never auto-crawls.
- New WebFeedConfig (parsers.webfeed): max_pages, concurrency, politeness
delay, same_host_only, respect_robots, max_depth, suggest_feed.
- Dependencies: feedparser (robust RSS/Atom), defusedxml (XXE-safe XML).
- Docs: zh/en resources API, MCP/CLI/SDK help, ov.conf.example.
- Tests: 52 unit tests (fake httpx, no network).
* chore: clear unused files
* fix(tests): fix unit test
* refactor(auth): introduce plugin-based authentication architecture
Replace the monolithic `openviking/server/auth.py` with an extensible
plugin-based auth system. This refactor extracts the three built-in modes
(`dev`, `api_key`, `trusted`) into separate `AuthPlugin` implementations,
adds a registry for third-party plugins, and preserves all existing behavior
while enabling custom authentication backends (e.g. LDAP, OIDC, mTLS).
Key changes:
- **New public API**: `AuthPlugin` (ABC) and `register_auth_plugin` decorator.
- **New registry**: `AuthPluginRegistry` supports runtime registration.
- **Built-in plugins**: `DevAuthPlugin`, `ApiKeyAuthPlugin`, `TrustedAuthPlugin`.
- **Config change**: `auth_mode` widened from `Literal` to `str` for custom modes.
- **Validation delegated**: `validate_server_config()` now delegates to the active
plugin's `validate_config()`, preserving existing validation semantics.
- **Router compatibility**: All existing `require_*` decorators and `resolve_identity`
/ `get_request_context` dependencies remain unchanged. Routers import the same
symbols from `openviking.server.auth`.
- **Tests**: `conftest.py` manually wires the DevAuthPlugin in ASGI tests (lifespan
not triggered). `test_auth.py` expanded with plugin registration and validation tests.
- **Docs**: `04-authentication.md` (en/zh) updated with plugin registration examples.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): fix trusted mode test
* fix(tests): fix unit test
* fix(cli): remove unexisted transaction observer
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* fix(skills): now we allow viking://agent/skills again, and optimize CLI for skills
* docs(skills): use -p instead of --parent in agent skills examples
Align the `ov skills add` examples in the context-types and viking-uri
docs with the short flag `-p` introduced for `ov skills list/find/show`,
so all four user-facing examples consistently demonstrate the short form
when targeting `viking://agent/skills`.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): error check for api key
* fix(tests): unit test wait until resource not busy
* fix(tests): unit test wait until resource not busy
* fix(sdk): args form in skills find
* fix(skills): pass target uri in request body
---------
Co-authored-by: claude-sonnet-4-6 <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* feat(fs): add count API for directory entry counting
Adds a dedicated `count` endpoint that returns the exact number of files
and sub-directories under a directory by traversing the filesystem,
distinct from `stat`'s vector-index-based estimate. Wired through
VikingFS, FSService, HTTP router and sync/async/local SDK clients.
* feat(cli): add `ov count` command for directory entry counting
Wires the new fs.count HTTP endpoint into the Rust CLI. Adds
`-r/--recursive` and `-a/--all` flags. Documentation updated with
CLI usage examples.
* fix
---------
Co-authored-by: dingben.db@bytedance.com <dingben.db@bytedance.com@bytedance.com>
* feat(ovpack): add v2 manifest and conflict policy
Add a portable OVPack manifest for scalar metadata and make imports validate scope, derived files, and conflicts before writing.
* fix(ovpack): remove import vectorize option
Make OVPack imports always rebuild vectors in the target environment, keep legacy packages compatible, and reject unsupported manifest versions before writing.
* fix(ovpack): remove force import alias
Use on_conflict as the single OVPack import conflict policy and reject removed force inputs.
* fix(ovpack): regenerate runtime vector metadata
Keep type portable but stop exporting or applying created_at, updated_at, and active_count from OVPack manifests.
* fix(ovpack): validate manifest contents
* fix(ovpack): require manifests for imports
* fix(ovpack): close manifest validation gaps
* fix(ovpack): defer parent creation until validation passes
* fix(ovpack): remove export size guard
* fix(ovpack): support session and scope-root restores
* docs(ovpack): document full backup migration
* feat(ovpack): add backup restore workflow
* fix(ovpack): validate import scope compatibility
* feat(server): add operation telemetry for session create/add_message/commit APIs
Wrap session.create, session.add_message and session.commit HTTP handlers with
run_operation so callers can opt in via TelemetryRequest and receive a
telemetry summary in the response. Propagate the telemetry parameter through
the async/sync HTTP clients, the local client and the public SDK so all
client modes expose a consistent surface.
* refactor(client/local): move part imports into _add_message_impl where they are used
Extend the ``target_uri`` parameter from ``str`` to ``Union[str, List[str]]``
across the full find/search stack so callers can scope a single query to
multiple directories in one request:
- server: FindRequest / SearchRequest accept list[str] target_uri
- service: SearchService.find / .search forward list[str]
- storage: VikingFS.find normalizes list[str], canonicalizes each entry,
and forwards the full list as target_directories to the retriever
(matching the existing behaviour of VikingFS.search)
- clients: BaseClient, LocalClient, AsyncHTTPClient, SyncHTTPClient,
AsyncOpenViking and SyncOpenViking signatures updated; the HTTP client
gains a ``_normalize_target_uri`` helper that applies
``VikingURI.normalize`` to each non-empty entry
Single-string behaviour is fully preserved: a plain ``str`` is normalized
internally to a one-element list, and empty ``""`` keeps today's
no-target semantics.
* feat(session): add account namespace policy and shared sessions
Unify namespace resolution across filesystem, indexing, and session storage.
Add account-shared session paths, role_id auth semantics, and an HTTP demo
script for the four namespace-policy combinations.
* space
* fix(pack): skip derived semantic files in ovpack transfer
Keep ovpack imports resilient to stale sidecars and rebuild semantics through the normal queue instead of restoring derived files verbatim.
* Revert "fix(pack): skip derived semantic files in ovpack transfer"
This reverts commit f4e4db8401.
* fix(namespace): default legacy accounts to agent-shared policy
Clarify that memory.agent_scope_mode is deprecated and document the supported agent memory migration paths.
* feat(sessions): support specifying session_id when creating session
* feat(session): validate session_id uniqueness on create
Add AlreadyExistsError check in session_service.create() when a specific
session_id is provided, ensuring idempotent behavior and preventing
accidental overwrites of existing sessions.
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
---------
Co-authored-by: openviking <openviking@example.com>
* feat(session): make commit two-phase with async memory extraction
Session commit now returns immediately after archiving messages (Phase 1).
Summary generation and memory extraction (Phase 2) run in the background
via asyncio.create_task(), returning a task_id for polling progress.
- Add get_task() API across all client layers for querying background task status
- get_session() auto-creates session if it does not exist
- Remove wait parameter and telemetry from commit endpoint
- Add .done completion marker to archive directories
- Update docs (EN/ZH) and tests for new two-phase flow
* feat(session): add .meta.json persistence and auto_create control for get_session
SessionService.get() now defaults to auto_create=False, raising NotFoundError
for missing sessions. A new SessionMeta dataclass tracks created_at, updated_at,
message_count, commit_count, memories_extracted (by category), last_commit_at,
and cumulative llm_token_usage. Meta is persisted to .meta.json and updated on
add_message, commit Phase 1 (message clear), and commit Phase 2 completion
(token usage, memory counts via bind_telemetry). All client layers
(local/async/sync/HTTP) and API docs updated accordingly.
* fix: remove session vectorize
* support commit for openclaw-plugin (#902)
Made-with: Cursor
* fix: reuse latest archive overview in session context
Thread the latest completed archive overview into archive summary generation and memory extraction, and simplify search context assembly to current messages plus the latest archive overview.
Co-Authored-By: Claude Opus 4.6
* refactor: session overview
---------
Co-authored-by: AutoCoder <wulf234@163.com>
The internal Session.exists() and SessionService.get() (added in #235)
provide session existence checks at the service layer, but these are
not accessible through the public client API (BaseClient and its
implementations).
This commit exposes two opt-in mechanisms at the public API level:
1. session(must_exist=True) — raises NotFoundError if the session
directory does not exist. Default must_exist=False preserves full
backward compatibility.
2. session_exists(session_id) — async convenience method returning
True/False without loading the full session.
Both delegate to the existing Session.exists() internally. No changes
to session.py or session_service.py — this is purely a public API
surface addition.
Files changed (7):
- base.py: updated abstract interface
- local.py: must_exist via Session.exists(), session_exists() delegate
- http.py: must_exist via get_session() HTTP call, session_exists()
- sync_http.py: pass-through
- async_client.py: session(must_exist) + session_exists()
- sync_client.py: pass-through
- test_session_lifecycle.py: 6 new tests
1. Complete parts support in all client layers:
- BaseClient (abstract interface)
- SyncHTTPClient
- AsyncOpenViking
- SyncOpenViking
2. Simplify VikingFS error handling:
- Remove _convert_agfs_error() complex error mapping
- Use simple FileNotFoundError for all AGFS exceptions
- This is cleaner and the original PR's error mapping was
over-engineered for the use case
* feat: add directory parsing support to OpenViking
- Implemented DirectoryParser to handle local directories with mixed document types.
- Enhanced add_resource function to support directory imports with options for including, excluding, and ignoring specific directories.
- Updated client and service layers to forward additional parsing options.
- Added unit tests for DirectoryParser to ensure correct functionality and error handling.
- Improved user feedback with rich table summaries for processed, failed, unsupported, and skipped files during directory imports.
* docs: update README.md to include directory import instructions for add.py
* style: reformat files to pass CI code formatting
* style: reformat files to pass CI code formatting