* fix(embedding): downsample oversized image inputs
Keep imported image resources unchanged while avoiding provider-side multimodal embedding failures for oversized images. The embedding path now builds a temporary downsampled image data URI when image bytes exceed the shared large-image limits.
Move reusable image size thresholds into media_limits so both parser-side large image handling and embedding-side input preparation depend on a common utility instead of embedding_utils importing parser internals.
Add vectorize_file coverage confirming large image embedding inputs are resized and the stored resource bytes are preserved.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(media): downsample oversized image model inputs
Keep imported image resources unchanged while avoiding provider-side multimodal failures for oversized images. Shared image input preparation now builds temporary downsampled bytes for model requests when image bytes exceed the configured large-image limits.
Apply the model-input downsampling to both semantic image summary generation and embedding image data URI construction, so directory and code repository imports can preserve original images while sending provider-compatible inputs.
Move reusable image size thresholds into media_limits so parser-side large image handling, VLM summary generation, and embedding preparation share common limits without embedding_utils importing parser internals.
Always convert downsampled model images to RGB before JPEG encoding so Pillow-openable modes such as LA or I;16 do not fall back to the original oversized bytes.
Add coverage confirming VLM image summaries, vectorize_file embedding inputs, and JPEG-incompatible image modes are resized while stored resource bytes are preserved.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: add audio and video understanding via VLM
* docs: design media resource guards
* fix: bound media staging concurrency
* fix: cap unknown-size media staging
* test: stage media in routing fake
* test: exercise media staging callbacks
* test: trim media understanding coverage
* chore: 清理实现计划文档
* fix: 修复多凭证切换问题
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
* fix(server): stop exporting raw query strings and buffering zip responses in observability
Sweep findings: B-03, B-13. Prevent query secrets from reaching traces and keep ZIP responses streaming.
(cherry picked from commit d8ac3dc33b)
* fix(session): tolerate missing/corrupt archive in Phase-2 replay
(NotFoundError / _ArchiveMessagesCorruptError) on a missing or corrupt
archive messages.jsonl instead of returning []. That PR added skip-on-
missing tolerance to the read path (_get_uncovered_archive_messages) and to
resume_queued_commit, but not to the Phase-2 commit replay path
(_prepare_phase2_archive_messages), which calls _read_archive_messages
unguarded while rolling earlier failed archives into the current commit.
Consequence: a terminally-failed earlier archive whose messages.jsonl is
missing/corrupt (legacy "no messages" terminal data, or produced by #3417's
own archive_read terminal path) makes every subsequent commit's Phase-2
extraction raise -> caught by _run_memory_extraction's except -> the current
archive is terminal-failed too. Because the poisoned archive is only removed
from replay once "covered" (which requires a later archive to complete), and
no later archive can ever complete, the session's memory extraction is
permanently poisoned. Raw messages are safe, but extraction is stuck.
Fix: wrap the replay-loop _read_archive_messages call in the same tolerance
_get_uncovered_archive_messages already uses -- skip + warn on not-found
(_is_storage_not_found) and on _ArchiveMessagesCorruptError, re-raise real
storage failures. The skipped archive stays in covered_failed so the current
archive's .done marks it covered, clearing the poison permanently.
Adds a regression test asserting the replay skips a failed archive with a
missing messages.jsonl (and marks it covered) instead of raising, and that a
real storage failure still propagates.
Follow-up to #3417.
(cherry picked from commit 5b8ec9e68a)
* fix(client): align client surfaces without leaking memory metadata
Reconstructs the client-parity work from upstream PR #3439 on current main and strips reserved memory metadata before line slicing in both embedded and HTTP reads.
Based-on: 48b411d58c
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(index): propagate semantic vectorization failures safely
Reconstructs upstream PR #3437 on current main, carries enqueue failures through SemanticDagExecutor, and drains the attempt's embedding tracker before retry-visible failure propagation.
Based-on: 02387deb09
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(core): close privacy and embedding failure gaps
* fix(memory): strip repeated metadata trailers
* fix(core): close public memory visibility gaps
* ci: skip embedding-dependent resource test without secrets
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(parse): distinguish mpegts from TypeScript ts
* fix(parse): tighten mpegts ts routing semantics
* fix(semantic): use file name for media summary type
---------
Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* refactor(parse): simplify resource ingestion routing
Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.
* fix(feishu): preserve sheet and bitable imports
Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.
* fix(parse): keep normalized Feishu content internal
Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.
* fix(feishu): parse bitable blocks embedded in sheets
Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.
* fix(feishu): download bitable attachment images
* refactor(parse): remove unused document converter
* refactor(parse): unify Understanding routing
* docs(parse): mark routing classification points
* docs(parse): complete wait routing flow
* fix(parse): preserve Feishu Base URL scope
* refactor(resource): separate ingestion submission from execution
* fix(resource): reject internal ingestion fields at public entry
* fix(task): recover add-resource jobs after restart
Persist asynchronous add-resource work in QueueFS so interrupted jobs can resume instead of leaving tasks running forever.
* fix(queue): omit parser args from prepared jobs
* fix(queue): fail when semantic source is missing
* fix(embedding): keep failover retryable when only some creds auth-fail (#2916)
Follow-up to #2919. classify_api_error scans an exception's whole message with
auth winning over transient, so an AllCredentialsFailedError whose message
contains both a credential's 401 and a later credential's transient 500/429 was
classified auth and (after #2919) dropped as terminal, even though retrying the
transient credential could still succeed.
Classify AllCredentialsFailedError from its structured per-credential classes
instead of the concatenated message: transient if any credential failed
transiently, quota if any hit quota, auth only when every credential auth-failed.
Addresses qin-ctx's review on #2919.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: simplify failover error aggregation
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(feishu): surface permission errors clearly and keep users on page
Map Feishu/Lark API failures to typed OpenViking errors with actionable hints, and keep Web Studio from treating HTTP 403 permission denials as session logout.
* fix(feishu): simplify API error mapping
* refactor(feishu): inline API error mapping
---------
Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* Refactor recursive web import into HTTP accessor
Move ordinary web page import routing into HTTPAccessor and materialize crawled pages as a temporary directory via WebImporter.
Relocate Scrapy/Playwright crawling under parse.accessors.web_crawler, keep trafilatura extraction inside HTMLParser, and avoid repeated ResourceService.add_resource calls.
Add recursive crawl controls, safe request validation, page/download classification, and focused unit coverage.
* Document recursive web crawler options
* fix(web-crawler): stop SSRF sub-resource block from failing whole render
The playwright fallback validated every sub-resource request against the
SSRF guard and raised on the first disallowed host, failing the entire
page render. volcengine docs load a probe resource on an internal host,
so rendering always failed and the crawler stored the static anti-bot
"Please wait..." challenge page as content.
Now a blocked sub-resource is only aborted; the main document and final
URL still gate the result. Also wait past JS interstitials, retry reads
through in-flight navigation, and reject shell/challenge pages instead of
storing them.
* fix(web-crawler): surface renderer error hint on entry-page failure
When Playwright is unavailable, the renderer returns an actionable install
hint via RenderResult.error, but the spider silently kept the static shell
and WebImporter raised only the generic "Failed to fetch entry page". The
hint never reached the user.
Now the spider records rendered.error on the failed page, and WebImporter
appends the entry page's failure reason to the raised message so the CLI
shows the Playwright install instructions.
* fix(web-crawler): surface render hints and enforce crawl limits
* fix(web-crawler): avoid rendering SSR app pages
* perf(web-crawler): bound render concurrency and cap networkidle wait
Playwright renders were dispatched from parse callbacks without any
concurrency limit, so a page with many child links could spawn dozens of
Chromium pages at once (observed peak 28 for a 20-page crawl), risking OOM
on large sites and starting ~2.3x more renders than needed before
max_pages stopped the crawl. Gate renders with a semaphore sized to
config.concurrency and re-check the success limit after acquiring a slot
so queued callbacks skip rendering once the crawl is already done.
Also cap the networkidle wait at 8s: pages with continuous background
activity (e.g. GraphiQL) never go idle and previously blocked until the
full render timeout, turning a ~3s page into ~38s. Content is ready after
domcontentloaded and _wait_past_challenge covers late-arriving text.
Bump default concurrency 5 -> 10.
* fix(web-crawler): route .html/.htm URLs through recursive WebImporter
An explicit .html/.htm URL is detected as DOWNLOAD_HTML via the extension
map, so access() previously only routed URLType.WEBPAGE to WebImporter and
these URLs fell through to single-file download, silently ignoring
depth/max_pages. Route DOWNLOAD_HTML through WebImporter too, treating a
single-page import as the depth=0 case.
* fix(web-crawler): improve HTML extraction and rendering heuristics
- Drop trafilatura favor_precision=True: it stripped the full body of
link-dense pages, keeping only headers.
- Only render __NEXT_DATA__ pages with Playwright when their static body
is too thin; SSR/SSG Next.js pages already ship full text.
- Disable Scrapy telnet console to avoid opening port 6023.
* fix(web-crawler): keep code-hosting single-file URLs off recursive crawler
GitHub/GitLab blob and GitHub raw URLs resolve to a single file, not a
site. Route them through the single-file download path instead of the
recursive WebImporter, which otherwise crawls the hosting UI shell.
* docs(resources): add recursive web crawler usage examples
Add depth/max_pages crawl examples to the HTTP, Python SDK, and CLI
blocks in both the zh and en resource API docs, plus path-prefix
filtering and skip_download_links variants.
* cli: warn when auto-naming creates numbered copy
Adds a warning to the response when import creates a numbered copy
(e.g., resource_1) because the target already exists. The warning
includes a tip about using --to to specify an explicit target path.
The warning is returned in result['warnings'] so the CLI can display it.
Addresses #2707 (partial - improves discoverability of --to flag)
Per maintainer feedback on #2777.
* fix: simplify auto-naming warning
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat: understanding api support no wait method
* feat: delete gc
* fix: make staging dir in shared mode
* fix: append watch manager
* fix: alignment router function
* feat: supplement memory and persistence
* feat: change extra lock
* fix: solve several problems
* feat: adapt telemetry
* feat: fix some problems
* chore: clear unused files
* fix(tests): fix unit test
* refactor(auth): introduce plugin-based authentication architecture
Replace the monolithic `openviking/server/auth.py` with an extensible
plugin-based auth system. This refactor extracts the three built-in modes
(`dev`, `api_key`, `trusted`) into separate `AuthPlugin` implementations,
adds a registry for third-party plugins, and preserves all existing behavior
while enabling custom authentication backends (e.g. LDAP, OIDC, mTLS).
Key changes:
- **New public API**: `AuthPlugin` (ABC) and `register_auth_plugin` decorator.
- **New registry**: `AuthPluginRegistry` supports runtime registration.
- **Built-in plugins**: `DevAuthPlugin`, `ApiKeyAuthPlugin`, `TrustedAuthPlugin`.
- **Config change**: `auth_mode` widened from `Literal` to `str` for custom modes.
- **Validation delegated**: `validate_server_config()` now delegates to the active
plugin's `validate_config()`, preserving existing validation semantics.
- **Router compatibility**: All existing `require_*` decorators and `resolve_identity`
/ `get_request_context` dependencies remain unchanged. Routers import the same
symbols from `openviking.server.auth`.
- **Tests**: `conftest.py` manually wires the DevAuthPlugin in ASGI tests (lifespan
not triggered). `test_auth.py` expanded with plugin registration and validation tests.
- **Docs**: `04-authentication.md` (en/zh) updated with plugin registration examples.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): fix trusted mode test
* fix(tests): fix unit test
* fix(cli): remove unexisted transaction observer
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* fix(skills): now we allow viking://agent/skills again, and optimize CLI for skills
* docs(skills): use -p instead of --parent in agent skills examples
Align the `ov skills add` examples in the context-types and viking-uri
docs with the short flag `-p` introduced for `ov skills list/find/show`,
so all four user-facing examples consistently demonstrate the short form
when targeting `viking://agent/skills`.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): error check for api key
* fix(tests): unit test wait until resource not busy
* fix(tests): unit test wait until resource not busy
* fix(sdk): args form in skills find
* fix(skills): pass target uri in request body
---------
Co-authored-by: claude-sonnet-4-6 <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
Extends #2774 to the resource indexing paths. #2774 capped the memory abstract
at 50KB before vector writes (memory_updater._truncate_memory_abstract) but the
shared resource paths (vectorize_directory_meta / vectorize_file, which
index_resource feeds) still wrote the `abstract` scalar uncapped. An
.abstract.md or generated summary exceeding 65535 UTF-8 bytes raises the
vector-store bytes_row limit ("string field 'abstract' exceeds 65535 bytes")
and fails embedding enqueue, so the resource is silently never vectorized (and
thus not retrievable).
Add a _truncate_abstract_bytes helper mirroring #2774's 50KB cap and apply it
to the abstract scalar on the resource write paths. Adds regression tests for
the helper and both vectorize paths.