* feat(agent-evolution): reload global switch at commit time
* feat(agent-evolution): expose configured account in status
* test(agent-evolution): cover account in status response
* fix(agent-evolution): align live config reload semantics
* fix(agent-evolution): tolerate non-object live config
* fix(usage-reporter): use snake case count fields
* feat(usage-reporter): add file log sink
* fix
* fix: address live reload and usage sink review findings
* fix(usage-reporter): complete file sink compatibility
* fix: make experience snapshot source unambiguous
* docs(usage-reporter): align count record implementation plan
* fix: address agent evolution review blockers
* fix(usage-reporter): preserve Windows rollover deadline
* fix(usage-reporter): encode file records as JSON envelopes
* fix(usage-reporter): use snake case unique id
* fix(parse): distinguish mpegts from TypeScript ts
* fix(parse): tighten mpegts ts routing semantics
* fix(semantic): use file name for media summary type
---------
Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
* refactor(parse): simplify resource ingestion routing
Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.
* fix(feishu): preserve sheet and bitable imports
Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.
* fix(parse): keep normalized Feishu content internal
Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.
* fix(feishu): parse bitable blocks embedded in sheets
Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.
* fix(feishu): download bitable attachment images
* refactor(parse): remove unused document converter
* refactor(parse): unify Understanding routing
* docs(parse): mark routing classification points
* docs(parse): complete wait routing flow
* fix(parse): preserve Feishu Base URL scope
* refactor(resource): separate ingestion submission from execution
* fix(resource): reject internal ingestion fields at public entry
* fix(kernel): stop conflating storage failures with not-found
Sweep findings: A-03, A-07, A-11, A-12, B-07. Preserve storage and parse failures instead of reporting missing or empty state.
* fix(review): restore archive failure handling
Addresses blocking review finding on #3417.
* fix(review): terminalize corrupt archive records
Addresses blocking review finding on #3417.
* test: adapt pending-archive-skip test to refactored archive scan
Rebase onto main (#3380 turn-aware retention) changed archive refs to carry
an archive_id; update the test mock's _list_archive_refs return so the missing
pending archive still routes through _get_uncovered_archive_messages and is
skipped (not raised).
* fix(task): recover add-resource jobs after restart
Persist asynchronous add-resource work in QueueFS so interrupted jobs can resume instead of leaving tasks running forever.
* fix(queue): omit parser args from prepared jobs
* fix(queue): fail when semantic source is missing
* feat(connector): delegate add_resource imports to external Connector
Opt-in integration that routes add_resource data fetching and parsing
to external Connector service; the Connector stages source data and
calls back into OV through the standard add_resource pipeline.
- add ConnectorClient wrapping the control plane's inner doc/add and
task/info endpoints
- add [connector] config section: enable, connector/tracker endpoint
URLs, timeout_seconds, poll_interval_ms, allowed_add_types
- route add_resource via Connector when enabled and args.add_type is
in allowed_add_types; otherwise fall back to the standard pipeline
with an info log
- track imports as connector_import TaskRecords and poll Connector
task status in the background until terminal state or timeout
* fix(connector): delegate add_resource imports to external Connector
* Refactor recursive web import into HTTP accessor
Move ordinary web page import routing into HTTPAccessor and materialize crawled pages as a temporary directory via WebImporter.
Relocate Scrapy/Playwright crawling under parse.accessors.web_crawler, keep trafilatura extraction inside HTMLParser, and avoid repeated ResourceService.add_resource calls.
Add recursive crawl controls, safe request validation, page/download classification, and focused unit coverage.
* Document recursive web crawler options
* fix(web-crawler): stop SSRF sub-resource block from failing whole render
The playwright fallback validated every sub-resource request against the
SSRF guard and raised on the first disallowed host, failing the entire
page render. volcengine docs load a probe resource on an internal host,
so rendering always failed and the crawler stored the static anti-bot
"Please wait..." challenge page as content.
Now a blocked sub-resource is only aborted; the main document and final
URL still gate the result. Also wait past JS interstitials, retry reads
through in-flight navigation, and reject shell/challenge pages instead of
storing them.
* fix(web-crawler): surface renderer error hint on entry-page failure
When Playwright is unavailable, the renderer returns an actionable install
hint via RenderResult.error, but the spider silently kept the static shell
and WebImporter raised only the generic "Failed to fetch entry page". The
hint never reached the user.
Now the spider records rendered.error on the failed page, and WebImporter
appends the entry page's failure reason to the raised message so the CLI
shows the Playwright install instructions.
* fix(web-crawler): surface render hints and enforce crawl limits
* fix(web-crawler): avoid rendering SSR app pages
* perf(web-crawler): bound render concurrency and cap networkidle wait
Playwright renders were dispatched from parse callbacks without any
concurrency limit, so a page with many child links could spawn dozens of
Chromium pages at once (observed peak 28 for a 20-page crawl), risking OOM
on large sites and starting ~2.3x more renders than needed before
max_pages stopped the crawl. Gate renders with a semaphore sized to
config.concurrency and re-check the success limit after acquiring a slot
so queued callbacks skip rendering once the crawl is already done.
Also cap the networkidle wait at 8s: pages with continuous background
activity (e.g. GraphiQL) never go idle and previously blocked until the
full render timeout, turning a ~3s page into ~38s. Content is ready after
domcontentloaded and _wait_past_challenge covers late-arriving text.
Bump default concurrency 5 -> 10.
* fix(web-crawler): route .html/.htm URLs through recursive WebImporter
An explicit .html/.htm URL is detected as DOWNLOAD_HTML via the extension
map, so access() previously only routed URLType.WEBPAGE to WebImporter and
these URLs fell through to single-file download, silently ignoring
depth/max_pages. Route DOWNLOAD_HTML through WebImporter too, treating a
single-page import as the depth=0 case.
* fix(web-crawler): improve HTML extraction and rendering heuristics
- Drop trafilatura favor_precision=True: it stripped the full body of
link-dense pages, keeping only headers.
- Only render __NEXT_DATA__ pages with Playwright when their static body
is too thin; SSR/SSG Next.js pages already ship full text.
- Disable Scrapy telnet console to avoid opening port 6023.
* fix(web-crawler): keep code-hosting single-file URLs off recursive crawler
GitHub/GitLab blob and GitHub raw URLs resolve to a single file, not a
site. Route them through the single-file download path instead of the
recursive WebImporter, which otherwise crawls the hosting UI shell.
* docs(resources): add recursive web crawler usage examples
Add depth/max_pages crawl examples to the HTTP, Python SDK, and CLI
blocks in both the zh and en resource API docs, plus path-prefix
filtering and skip_download_links variants.
* fix(web-studio): lazy-create playground sessions to prevent orphans
The AgentPanel used to call POST /sessions on every page load (via a
useEffect gated on botHealth.isSuccess), creating empty UUID sessions
even when the user never sent a message. These orphan sessions cluttered
the session list.
Fix: generate a client-side UUID as the initial sessionId instead of an
empty string. The session is lazily created on the backend by the first
addMessage call (POST /sessions/{id}/messages uses auto_create=True).
- Remove startSession() and its auto-create useEffect
- Always have a sessionId (createRandomUuid fallback)
- Simplify useChat params (persistMessages always true, no preview fallback)
- Remove dead isCreating UI branch
- Call onSessionChange on mount so the URL stays in sync
- Error retry button now uses handleNewSession()
Closes: orphan empty sessions created on playground page load
* feat(sessions): sort session list by filesystem modTime
- Backend: add mod_time to session list API response (from viking_fs.ls stat)
- Frontend: useSessionListByRecency sorts by mod_time instead of N detail fetches
- Frontend: sidebar and playground history use recency-sorted list
- Frontend: sortTreeEntries sorts directories by modTime descending
Reduces session list load from 1+N requests to 1 request.
* feat: understanding api support no wait method
* feat: delete gc
* fix: make staging dir in shared mode
* fix: append watch manager
* fix: alignment router function
* feat: supplement memory and persistence
* feat: change extra lock
* fix: solve several problems
* feat: adapt telemetry
* feat: fix some problems
* chore: clear unused files
* fix(tests): fix unit test
* refactor(auth): introduce plugin-based authentication architecture
Replace the monolithic `openviking/server/auth.py` with an extensible
plugin-based auth system. This refactor extracts the three built-in modes
(`dev`, `api_key`, `trusted`) into separate `AuthPlugin` implementations,
adds a registry for third-party plugins, and preserves all existing behavior
while enabling custom authentication backends (e.g. LDAP, OIDC, mTLS).
Key changes:
- **New public API**: `AuthPlugin` (ABC) and `register_auth_plugin` decorator.
- **New registry**: `AuthPluginRegistry` supports runtime registration.
- **Built-in plugins**: `DevAuthPlugin`, `ApiKeyAuthPlugin`, `TrustedAuthPlugin`.
- **Config change**: `auth_mode` widened from `Literal` to `str` for custom modes.
- **Validation delegated**: `validate_server_config()` now delegates to the active
plugin's `validate_config()`, preserving existing validation semantics.
- **Router compatibility**: All existing `require_*` decorators and `resolve_identity`
/ `get_request_context` dependencies remain unchanged. Routers import the same
symbols from `openviking.server.auth`.
- **Tests**: `conftest.py` manually wires the DevAuthPlugin in ASGI tests (lifespan
not triggered). `test_auth.py` expanded with plugin registration and validation tests.
- **Docs**: `04-authentication.md` (en/zh) updated with plugin registration examples.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): fix trusted mode test
* fix(tests): fix unit test
* fix(cli): remove unexisted transaction observer
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* fix(skills): now we allow viking://agent/skills again, and optimize CLI for skills
* docs(skills): use -p instead of --parent in agent skills examples
Align the `ov skills add` examples in the context-types and viking-uri
docs with the short flag `-p` introduced for `ov skills list/find/show`,
so all four user-facing examples consistently demonstrate the short form
when targeting `viking://agent/skills`.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): error check for api key
* fix(tests): unit test wait until resource not busy
* fix(tests): unit test wait until resource not busy
* fix(sdk): args form in skills find
* fix(skills): pass target uri in request body
---------
Co-authored-by: claude-sonnet-4-6 <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* fix(queuefs): skip semantic generation for non-directory memory URIs
A memory file reindexed with mode=semantic_and_vectors enqueues a
context_type="memory" semantic message whose URI is a file. on_dequeue
routes it to _process_memory_directory, which ls()'d the file, raised, and
the outer handler re-enqueued it as a transient error — forever, starving
the semantic queue. The entry is AGFS-persisted, so it survives a restart
and blocks `reindex --wait` and memory writes that wait on the queue.
_process_memory_directory now stat()s the URI first: a confirmed
non-directory (or missing) URI is marked done and skipped instead of
listed, so the message acks and the queue drains. When stat is unavailable
it falls through to the existing ls() path, leaving current behavior
unchanged.
Fixes#2734
* fix(queuefs): preserve retries for stat failures
* fix(reindex): skip memory semantics for file targets
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>