Commit Graph
162 Commits
Author SHA1 Message Date
Zonas ZhouandClaude 6e77291265 feat(pdf): refactor MinerU parsing to the official file_parse API (#3953)
* feat(pdf): refactor MinerU parsing to the official file_parse API

* feat(pdf): remove mineru_api_key from configuration and examples

* feat(pdf): preflight MinerU /health during service initialization

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-17 13:48:49 +08:00
MaojiaShengandTRAE CLI 5aed7f72b4 fix(media): downsample oversized image model inputs (#3965)
* fix(embedding): downsample oversized image inputs

Keep imported image resources unchanged while avoiding provider-side multimodal embedding failures for oversized images. The embedding path now builds a temporary downsampled image data URI when image bytes exceed the shared large-image limits.

Move reusable image size thresholds into media_limits so both parser-side large image handling and embedding-side input preparation depend on a common utility instead of embedding_utils importing parser internals.

Add vectorize_file coverage confirming large image embedding inputs are resized and the stored resource bytes are preserved.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(media): downsample oversized image model inputs

Keep imported image resources unchanged while avoiding provider-side multimodal failures for oversized images. Shared image input preparation now builds temporary downsampled bytes for model requests when image bytes exceed the configured large-image limits.

Apply the model-input downsampling to both semantic image summary generation and embedding image data URI construction, so directory and code repository imports can preserve original images while sending provider-compatible inputs.

Move reusable image size thresholds into media_limits so parser-side large image handling, VLM summary generation, and embedding preparation share common limits without embedding_utils importing parser internals.

Always convert downsampled model images to RGB before JPEG encoding so Pillow-openable modes such as LA or I;16 do not fall back to the original oversized bytes.

Add coverage confirming VLM image summaries, vectorize_file embedding inputs, and JPEG-incompatible image modes are resized while stored resource bytes are preserved.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-12 22:24:57 +08:00
c1345a1f7e feat(feishu):Support Feishu Drive folder and file imports (#3937)
* Support Lark drive folder and file URLs

* Handle partial Feishu folder import failures

* Fix Feishu drive folder path names

* Tighten Feishu drive folder tests

* test(feishu): consolidate drive import coverage

---------

Co-authored-by: haoxingjun <haoxingjun@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-12 16:01:00 +08:00
Zayn Jarvis dcca29364c fix(parse): stop silently dropping markdown YAML frontmatter (#3929)
Frontmatter was parsed into ParseResult.meta and removed from the body, but
that metadata is never persisted, so every ingested markdown file lost its
frontmatter with no way to read the fields back.

Parse frontmatter into meta unconditionally (it still drives doc_title) and
only remove it from the stored body when explicitly configured; that removal
is now off by default.
2026-08-11 14:14:53 +08:00
zihengli cd55ec89a6 feat(assets): support fixed Git commits, explicit targets, and private repository auth (#3703)
* feat/openviking_assets_support_git_commit_id

* feat/openviking_assets_support_to

* feat/private_git_support_watch

* fix: doc_and_ut

* fix: doc_and_ut

* fix: adapt git token url
2026-08-11 14:07:21 +08:00
Qin Haojie 7abd6ab249 refactor(client): remove Python embedded mode (#3712)
* refactor(client): remove Python embedded mode

Consolidate Python consumers on the HTTP SDK while keeping shared server and storage capabilities unchanged.

* refactor(client): remove obsolete embedded leftovers
2026-08-10 18:00:00 +08:00
bianbiandashen 3087f943a2 fix(markdown): keep force-split chunks within the token budget (#3672)
_smart_split_content documents that it enforces both a token limit
(max_size) and a hard character limit. But when a single paragraph was
oversized by tokens yet under the character limit, the force-split loop
stepped through it by max_chars only, so it emitted a chunk that still
exceeded max_size tokens.

This is reachable with ordinary long-form CJK text: _estimate_token_count
weights CJK at ~0.7 token/char, so a ~5000-char Chinese paragraph is
~3500 tokens (over the 2048 default) while staying under the char limit,
and was returned as a single over-budget chunk.

Bound the force-split step by min(max_chars, max_size / MAX_TOKENS_PER_CHAR),
where MAX_TOKENS_PER_CHAR is the worst-case (CJK) density already used by
_estimate_token_count, now extracted into a shared constant so the two stay
in sync.

Add regression tests for the token budget and content preservation.
2026-08-08 01:54:06 +08:00
bianbiandashen 7b8b33e8f8 fix(markdown): match GitHub anchor slugs for headings with punctuation (#3673)
_gh_slug claims to produce "GitHub-style" heading slugs but collapsed runs
of whitespace (`re.sub(r"\s+", "-", s)`). GitHub's reference slugger
(github-slugger) maps each space to its own hyphen (`.replace(/ /g, '-')`)
and does not collapse.

Because punctuation is stripped before spaces are converted, a heading like
"Foo & Bar" leaves two adjacent spaces where "&" was. GitHub renders this as
"foo--bar", but _gh_slug produced "foo-bar". The intra-document link rewriter
(_rewrite_link) compares _gh_slug(heading) against the link fragment, so an
author-written link such as `guide.md#foo--bar` failed to match its heading
and was left unrewritten after the target doc was split into sections.

Replace `\s+` with `\s` so each whitespace character maps to one hyphen,
matching GitHub. Simple single-space headings are unaffected.

Add regression tests covering punctuation headings and ordinary headings.
2026-08-07 20:56:40 +08:00
Kchen 8d1d52fe5d 资源导入:支持解析后不拆分文档 (#3645) 2026-08-05 11:34:10 +08:00
Haoyu ZhangandQin Haojie 3f3554256b feat: 支持基于火山方舟的音视频多模态理解 (#3563)
* feat: add audio and video understanding via VLM

* docs: design media resource guards

* fix: bound media staging concurrency

* fix: cap unknown-size media staging

* test: stage media in routing fake

* test: exercise media staging callbacks

* test: trim media understanding coverage

* chore: 清理实现计划文档

* fix: 修复多凭证切换问题

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-08-03 16:44:22 +08:00
yangxinxin-7andClaude Opus 5 8b4deaab99 fix(parser): drop duplicate images when a PDF stacks XObjects on one spot (#3662)
Print-to-PDF producers routinely emit several image XObjects drawn at the
exact same position on a page (a background layer plus a content layer).
Because `_extract_image_from_page` rasterises the page *region* rather than
decoding the XObject itself, every one of them renders to identical bytes —
so a document with two stacked full-page layers wrote two byte-identical
PNGs per page and referenced both from the generated markdown.

Dedup within each page, in two steps:

- bbox first, so a repeat is skipped before paying for the render;
- a content hash as a backstop, for bboxes that differ slightly but still
  rasterise to the same bytes.

Both sets are per-page, so a header logo repeated across pages is still
kept once on every page. `meta["images_deduplicated"]` reports how many
were skipped.

Measured on an 8-page article exported from a web page: 16 saved PNGs -> 8,
16 markdown image references -> 8, local conversion 3.8s -> 2.4s.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:03:21 +08:00
zgy c241a2a043 chore: reduce import log noise (#3657) 2026-07-31 15:31:44 +08:00
zgy 49b182045b refactor(parser): Refactor code summaries to fixed skeleton-first routing (#3568)
* Refactor code summary skeleton routing

* Simplify code skeleton routing configuration

* Render C tag skeletons as signatures

* Revert "Render C tag skeletons as signatures"

This reverts commit 8e342055f8.

* Simplify fixed code skeleton summary route

* Inline process skeleton rendering

* Simplify code skeleton routing entrypoints

* Fix code summary review issues

* Address final code summary review feedback

* Route failed tags skeletons to LLM fallback

* Restore CUDA and TS extension routing

* Improve code skeleton query coverage

* Route semantic code detection through skeleton support

* Move process skeleton engine into ast package

* Admit skeleton-supported files during directory scan

* Align code summary docs after main merge

* Reduce code skeleton fallback log verbosity

* chore: require grep-ast 0.9.0
2026-07-31 11:38:57 +08:00
Eurakaxun 44c6df2622 perf: retrieval, import, LangChain, and session-context optimizations (#3569) 2026-07-30 10:10:11 +08:00
baojun-zhang 2f9451231e refactor(pathlock):using rust implement instead python (#3602)
* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):optimize unit test code

* refactor(pathlock):optimize encryption create func

* refactor(pathlock):avoid releasing handoffed pathlock on enqueue errors

* fix(pathlock): use owned lease capability and handle S3 create-new 409 as conflict

* fix(ragfs): keep original FsContext for multi-write metadata

* fix(pathlock): resolve lease coverage and CAS handling issues

- detect S3 conditional conflicts from structured service errors
- pass transaction leases when deleting skill roots
- let temp cleanup acquire locks for temp paths
- disambiguate cache and pathlock providers in cache tests
- update temp cleanup lease assertions

* fix(ragfs): bypass pathlock for multi-write metadata

* fix(ragfs): revert pathlock fail-fast design

* fix(ragfs):fix(ragfs): use non-blocking fcntl locks for localfs CAS

* fix(ragfs): serialize heartbeat lease refresh with release and report real conflict kind

* fix(ragfs): preserve conflict kind snapshot and drop unused test scaffolding

* fix(ragfs): preserve conflict kind snapshot and drop unused test scaffolding
2026-07-29 19:45:34 +08:00
chenxiaobin-monkeyandchenxiaobin.monkey ff37e25cfd fix(parse): distinguish mpegts from TypeScript ts (#3574)
* fix(parse): distinguish mpegts from TypeScript ts

* fix(parse): tighten mpegts ts routing semantics

* fix(semantic): use file name for media summary type

---------

Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
2026-07-29 13:28:25 +08:00
baojun-zhang 1841dfed81 Revert "refactor(pathlock):using rust implement instead python (#3557)" (#3597)
This reverts commit 6b538db569.
2026-07-29 11:31:41 +08:00
baojun-zhang 6b538db569 refactor(pathlock):using rust implement instead python (#3557)
* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):optimize unit test code

* refactor(pathlock):optimize encryption create func

* refactor(pathlock):avoid releasing handoffed pathlock on enqueue errors

* fix(pathlock): use owned lease capability and handle S3 create-new 409 as conflict

* fix(ragfs): keep original FsContext for multi-write metadata

* fix(pathlock): resolve lease coverage and CAS handling issues

- detect S3 conditional conflicts from structured service errors
- pass transaction leases when deleting skill roots
- let temp cleanup acquire locks for temp paths
- disambiguate cache and pathlock providers in cache tests
- update temp cleanup lease assertions

* fix(ragfs): bypass pathlock for multi-write metadata

* fix(ragfs): revert pathlock fail-fast design
2026-07-29 11:08:42 +08:00
zihengli a1e468b982 feat(connector): support more git like platform (#3531)
* feat(connector): support more git like platform

* feat(connector): support more git like platform

* feat(connector): support more git like platform

* feat(connector): support more git like platform

* feat(connector): support more git like platform
2026-07-28 19:17:47 +08:00
Evo 16e7c33f1f fix(feishu): preserve Bitable media permission context during import (#3558) 2026-07-28 11:41:39 +08:00
Qin Haojie 2be4bb4879 refactor(parse): 收口资源解析路由 (#3295)
* refactor(parse): simplify resource ingestion routing

Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.

* fix(feishu): preserve sheet and bitable imports

Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.

* fix(parse): keep normalized Feishu content internal

Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.

* fix(feishu): parse bitable blocks embedded in sheets

Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.

* fix(feishu): download bitable attachment images

* refactor(parse): remove unused document converter

* refactor(parse): unify Understanding routing

* docs(parse): mark routing classification points

* docs(parse): complete wait routing flow

* fix(parse): preserve Feishu Base URL scope

* refactor(resource): separate ingestion submission from execution

* fix(resource): reject internal ingestion fields at public entry
2026-07-27 16:47:20 +08:00
Wu JiaCheng 5ab30d3e04 fix(parse): preserve text file parser metadata (#3480) 2026-07-24 16:30:38 +08:00
Wu JiaChengandqin-ctx 62a913795c fix(parse): honor markdown frontmatter config (#3475)
* fix(parse): honor markdown frontmatter config

* refactor(parse): simplify frontmatter config handling

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-23 11:55:59 +08:00
d681ef3158 fix(parse): handle parentheses in Markdown image paths (#3462)
* fix(parse): handle parentheses in Markdown image paths

The image regex !\[([^\]]*)\]\(([^)]+)\) used [^)]+ for the path capture
group, which truncates at the first ) character. When document titles
or filenames contain balanced parentheses (e.g. "文档_17 (17号项目)"), the
generated image paths include ) and the regex captures a truncated,
non-existent path. This causes _resolve_image_path() to fail silently
(WARNING only), and the image is never copied to VikingFS or sent to
VLM for understanding.

Fix: replace the path capture group with (?:[^()]|\([^()]*\))+, which
allows one level of balanced parentheses inside the path while still
terminating at the correct closing ) of the Markdown image syntax.

Add focused tests covering balanced parens in directory and filename
components, URLs with parens, multiple images on one line, and
non-matching of plain links.

Fixes #3455

* fix(test): exercise MarkdownParser._image_pattern directly, remove unused import

Address review feedback on #3462:
1. Tests now import and instantiate MarkdownParser to access the
   production _image_pattern regex, instead of compiling an independent
   copy. Tests fail if the production regex regresses.
2. Remove unused `import pytest` (Ruff F401).

* fix(parse): rewrite parenthesized image paths

* test: remove extra image rewrite regression case

* fix(parse): share markdown image parsing for rewrite

* refactor(parse): keep markdown image fix minimal

---------

Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-23 11:33:26 +08:00
Qin Haojie fd098cfd65 fix(cli): preserve structured API errors (#3379)
* fix(error): preserve structured errors in cli

* fix(cli): preserve status for non-json errors
2026-07-22 18:06:48 +08:00
Wu JiaCheng 061359a2f5 fix(parse): normalize MIME aliases with parameters (#3393) 2026-07-22 15:44:43 +08:00
Wu JiaCheng 63c878e60e fix(parse): import extensionless README files (#3394) 2026-07-22 15:29:13 +08:00
ShaoZegangByte 39dc01e2a5 feat(resource): support Feishu/Lark URL imports via UnderstandingAPI (#3320)
* feat(resource):add lark understand api

* fix(resource):add lark understand api env

* fix(resource):add lark understand api

* fix(resource):understand api pr

* fix(resource):add test

* fix(resource):handle deferred URIs for async Feishu imports

* fix(resource):fix Ruff issues in lark import

* fix: persist final resource URI for async UnderstandingAPI tasks

* fix: clean up cancelled async UnderstandingAPI scheduling

* fix: getattr defer_target_resolution
2026-07-21 11:08:13 +08:00
huangruitengandhuangruiteng f093fafd1b fix(feishu): preserve title prefixes in resource names (#3366)
Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-20 17:16:49 +08:00
MaojiaShengandqin-ctx 3c70f8d370 feat(parse): add large image processing for image parser (#3265)
* feat(parse): add large image processing for image parser

- Add large_image_processor.py: detect large images (>10MB or >4096px),
  create low-res previews, split into grid tiles, and generate grid
  overlay images with tile labels
- Refactor ImageParser.parse() to integrate large image processing pipeline
- Enable SVG-to-PNG conversion in utils.py (cairosvg/wand)
- Rename ImageConfig.max_dimension to preview_max_dimension and add new
  config fields: max_file_size_mb, max_tile_size_mb, max_tile_dimension_px,
  tile_overlap_px, large_image_threshold_dimension
- Update ov.conf.example with new image config options

* fix(parse): correct tile dimension comment from 1024px to 2048px

* fix(parse): fix tile label path in grid overlay to include tiles/ directory

* fix(parse): register missing image extensions for ImageParser

TIFF, ICO, DIB, ICNS, SGI, JP2 were not in IMAGE_EXTENSIONS, causing
them to fallback to TextParser. All are supported by PIL.

* fix(parse): preserve PNG format for tiles instead of always converting to JPEG

* fix(parse): address review feedback for large image processing

- Wire config.image to ImageParser in ParserRegistry (was missing)
- Remove unnecessary preview creation for small images (broke LA mode PNG)
- Enforce max_tile_size_mb on tiles with quality reduction and resize fallback
- Remove 64-tile hard cap that conflicted with max_tile_dimension_px
- Add comment explaining why original file is not saved for large images

* refactor(parse): remove max_tile_size_mb as it is a soft suggestion

max_tile_size_mb was a soft constraint that was not enforced
consistently. Remove it from config, constants, and all enforcement
logic. Tile dimension (max_tile_dimension_px) remains the sole constraint.

* fix(parse): use CJK-capable font for grid overlay labels

The old font loading only tried macOS-specific paths and fell back to
PIL's default bitmap font, which cannot render CJK characters in
filenames. Add a cross-platform CJK font lookup that covers Linux
(Noto/Droid/WQY/DejaVu), macOS (PingFang), and Windows (MSYH/SimSun).

* fix(parse): convert non-VLM-supported image formats to PNG on save

Image formats like TIFF, ICO, DIB, ICNS, SGI, JP2 are not recognized
by VLM backends (OpenAI/LiteLLM/VolcEngine only support PNG/JPEG/GIF/
WebP/BMP) or by embedding_utils for image vectorization. When a file
with one of these extensions is parsed, convert it to PNG and use a
.png extension so that downstream pipelines see consistent data.

SVG files (already PNG-converted via cairosvg) also get the .png
extension for the same reason.

* fix(parse): import io for SVG conversion

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-16 17:37:05 +08:00
Qin Haojie d47f2106ee refactor: remove unused and deprecated APIs (#3272)
Delete dead compatibility paths and test-only helpers so unsupported APIs do not remain as accidental contracts.
2026-07-16 10:49:56 +08:00
Haoyu Zhangandzhanghaoyu c6d48bc056 feat(parse): support AC-3 audio resources and add multimodal integration tests (#3229)
* docs: design AC-3 resource ingestion

* chore: ignore local worktrees

* feat(parse): support AC-3 audio resources

* feat: 补充多模态文档解析测试脚本

---------

Co-authored-by: zhanghaoyu <zhanghaoyu.la@bytedance.com>
2026-07-14 11:03:49 +08:00
chenxiaobin-monkeyandchenxiaobin.monkey b808aa0791 feat: 添加ts格式 (#3209)
Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
2026-07-13 13:55:44 +08:00
huangruitengandhuangruiteng cbcec52d7d fix(parse): preserve filters for local git repositories (#3190)
Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-13 11:35:21 +08:00
Hao Zhe 47170b05dc fix(parse): route mislabeled OOXML Word files correctly (#3113) 2026-07-10 14:29:59 +08:00
Qin Haojie cbdab57353 fix(security): prevent Git submodule SSRF (#3108)
Disable recursive submodule fetching and avoid exposing remote Git errors to resource API callers.
2026-07-10 11:55:19 +08:00
Kchenandchenpengfei e19d5f7d66 修复签名视频链接的解析路由 / Fix signed video URL parser routing (#3103)
中文:对 HTTP/HTTPS URL 使用 urlparse(url).path 提取扩展名,确保 video.mp4?signature=... 命中 UnderstandingAPI fast path。新增带签名视频 URL 的 ParserRouter 回归测试。

English: Parse HTTP/HTTPS URL paths before checking extensions so signed video URLs hit the UnderstandingAPI fast path. Add a ParserRouter regression test for signed video URLs.

Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
2026-07-09 20:34:20 +08:00
zgy 07113f81e0 fix(cli): don't misreport remote fetch auth failures as API key errors (#3074)
When the server crawls a URL that returns 401/403, it wraps the failure
in a 5xx envelope. The CLI matched on auth-flavored message text alone
and rendered "OpenViking rejected the API key", wrongly pointing users
at their local config.

Gate the API-key error report on status: 5xx responses are never treated
as a client auth failure even when the message mentions authentication or
forbidden. Also split the 401 vs 403 fetch messages so 403 reads as an
access-denied / anti-bot block rather than a credential problem.
2026-07-08 14:50:59 +08:00
cd9add7a27 fix(feishu): surface permission errors clearly and keep users on page (#3032)
* fix(feishu): surface permission errors clearly and keep users on page

Map Feishu/Lark API failures to typed OpenViking errors with actionable hints, and keep Web Studio from treating HTTP 403 permission denials as session logout.

* fix(feishu): simplify API error mapping

* refactor(feishu): inline API error mapping

---------

Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-07 19:55:43 +08:00
07aa9dc775 feat(feishu): persist imported document images (#3033)
* feat(feishu): persist imported document images

Download Feishu image tokens into local import temp trees so Markdown image references can be ingested alongside the document.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(feishu): harden inline-image download (async, extension, user token)

Address review feedback on the inline-image import path:

- Run the synchronous lark-oapi media download via asyncio.to_thread so a
  slow Feishu request no longer blocks unrelated async work on the event loop,
  matching the existing _fetch_document() pattern.
- Infer the image file extension from the downloaded bytes (byte-magic
  sniffing) and fall back to the response Content-Type via the existing
  mime_types.get_preferred_extension helper, instead of hardcoding .png. This
  stops JPEG/WebP/GIF bytes from being mislabeled as PNG to downstream
  consumers (e.g. the data:image/... URI built during multimodal vectorization).
- Advertise AccessTokenType.USER on the media download request when a user
  access token is supplied, so lark-oapi actually injects it. Previously the
  request only allowed TENANT, so user-token imports read the document body
  but silently dropped every image.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 17:44:26 +08:00
baojun-zhang 6a33ebb7ca Optimize glob walkdir (#3013)
* feat(storage): optimize glob func

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* fix(localfs): offload blocking fs operations to spawn_blocking

* feat(glob): cap glob api default node_limit at 256

* feat(sdk): add node_limit options for glob in python and go SDKs
2026-07-06 21:39:16 +08:00
zgy 8d861fabfe fix(web-crawler): remove Playwright rendering andoptimize the robots.txt entry-failure message (#3040)
* refactor(web-crawler): remove Playwright rendering, keep static SPA shells

Drop the Playwright-based fallback rendering path so SPA pages are stored
as their static shell HTML instead of being rendered headlessly. SPA
shells now surface the <noscript> notice (e.g. "You need to enable
JavaScript to run this app.") rather than producing an empty document,
and robots.txt-blocked entry pages get a human-readable error message.

- Delete playwright_renderer.py, render_heuristics.py and their tests
- Strip fallback_playwright/playwright_timeout config, fallback_rendered
  counter, and render-hint plumbing from the crawler and web importer
- HTMLParser: drop the SPA-empty-pattern stripping and fall back to
  <noscript> text when trafilatura extracts nothing
- Humanize the robots.txt entry-failure message
- Migrate the orphaned _convert_to_raw_url tests to HTTPAccessor (the
  method moved there in an earlier reorg) and remove the stale file

* refactor(web-importer): simplify robots.txt-blocked import message

Replace the verbose robots.txt explanation with a short compliance hint
that omits the URL and points users to local-file import instead.
2026-07-06 18:48:39 +08:00
MaojiaSheng 47da6ce129 refactor(bot): simplify vikingbot installation - merge all bot-* extras into [bot] (#3037) 2026-07-06 16:17:19 +08:00
zgy a50e9fd677 feat: add recursive web crawler based on Scrapy (#2836)
* Refactor recursive web import into HTTP accessor

Move ordinary web page import routing into HTTPAccessor and materialize crawled pages as a temporary directory via WebImporter.

Relocate Scrapy/Playwright crawling under parse.accessors.web_crawler, keep trafilatura extraction inside HTMLParser, and avoid repeated ResourceService.add_resource calls.

Add recursive crawl controls, safe request validation, page/download classification, and focused unit coverage.

* Document recursive web crawler options

* fix(web-crawler): stop SSRF sub-resource block from failing whole render

The playwright fallback validated every sub-resource request against the
SSRF guard and raised on the first disallowed host, failing the entire
page render. volcengine docs load a probe resource on an internal host,
so rendering always failed and the crawler stored the static anti-bot
"Please wait..." challenge page as content.

Now a blocked sub-resource is only aborted; the main document and final
URL still gate the result. Also wait past JS interstitials, retry reads
through in-flight navigation, and reject shell/challenge pages instead of
storing them.

* fix(web-crawler): surface renderer error hint on entry-page failure

When Playwright is unavailable, the renderer returns an actionable install
hint via RenderResult.error, but the spider silently kept the static shell
and WebImporter raised only the generic "Failed to fetch entry page". The
hint never reached the user.

Now the spider records rendered.error on the failed page, and WebImporter
appends the entry page's failure reason to the raised message so the CLI
shows the Playwright install instructions.

* fix(web-crawler): surface render hints and enforce crawl limits

* fix(web-crawler): avoid rendering SSR app pages

* perf(web-crawler): bound render concurrency and cap networkidle wait

Playwright renders were dispatched from parse callbacks without any
concurrency limit, so a page with many child links could spawn dozens of
Chromium pages at once (observed peak 28 for a 20-page crawl), risking OOM
on large sites and starting ~2.3x more renders than needed before
max_pages stopped the crawl. Gate renders with a semaphore sized to
config.concurrency and re-check the success limit after acquiring a slot
so queued callbacks skip rendering once the crawl is already done.

Also cap the networkidle wait at 8s: pages with continuous background
activity (e.g. GraphiQL) never go idle and previously blocked until the
full render timeout, turning a ~3s page into ~38s. Content is ready after
domcontentloaded and _wait_past_challenge covers late-arriving text.

Bump default concurrency 5 -> 10.

* fix(web-crawler): route .html/.htm URLs through recursive WebImporter

An explicit .html/.htm URL is detected as DOWNLOAD_HTML via the extension
map, so access() previously only routed URLType.WEBPAGE to WebImporter and
these URLs fell through to single-file download, silently ignoring
depth/max_pages. Route DOWNLOAD_HTML through WebImporter too, treating a
single-page import as the depth=0 case.

* fix(web-crawler): improve HTML extraction and rendering heuristics

- Drop trafilatura favor_precision=True: it stripped the full body of
  link-dense pages, keeping only headers.
- Only render __NEXT_DATA__ pages with Playwright when their static body
  is too thin; SSR/SSG Next.js pages already ship full text.
- Disable Scrapy telnet console to avoid opening port 6023.

* fix(web-crawler): keep code-hosting single-file URLs off recursive crawler

GitHub/GitLab blob and GitHub raw URLs resolve to a single file, not a
site. Route them through the single-file download path instead of the
recursive WebImporter, which otherwise crawls the hosting UI shell.

* docs(resources): add recursive web crawler usage examples

Add depth/max_pages crawl examples to the HTTP, Python SDK, and CLI
blocks in both the zh and en resource API docs, plus path-prefix
filtering and skip_download_links variants.
2026-07-03 19:25:10 +08:00
baojun-zhang 56c7f3d472 fix(queuefs): reuse outer lock_handle for sidecar, image rewrite and experience writebacks (#2953) 2026-07-02 16:48:38 +08:00
t0saki 2846bb6e76 feat(resources): ingest whole sites via sitemap / RSS / Atom (#2858)!
Add WebFeedAccessor (priority 60) that turns a single sitemap /
sitemapindex / RSS / Atom URL into ONE resource tree: it mirrors every
listed page into a temp directory and reuses the existing DirectoryParser
pipeline (the same "fetch-many -> dir -> tree" contract as GitAccessor).
A watch on the feed URL keeps the whole site refreshed (new pages added,
removed pages dropped on each rebuild).

- New openviking/parse/accessors/web_feed_accessor.py: WebFeedAccessor +
  sitemap/feed extractors (nested sitemapindex recursion with depth cap,
  RSS 2.0 / Atom via feedparser), bounded concurrent polite mirroring,
  robots.txt, same-host / include / exclude / max_pages limits.
- args={"site": true} forces whole-site ingestion from a bare domain or
  page by auto-discovering the sitemap/RSS (robots.txt, HTML
  <link rel=alternate>, conventional paths); {"site": false} opts a
  feed-looking URL back out to HTTPAccessor.
- Thread accessor-selection kwargs through can_handle; the registry
  tolerates accessors whose can_handle lacks **kwargs (back-compatible).
- Single-page adds get a non-blocking "this site exposes a sitemap/RSS"
  suggestion appended to the MCP add_resource response, gated to the
  site root only; never auto-crawls.
- New WebFeedConfig (parsers.webfeed): max_pages, concurrency, politeness
  delay, same_host_only, respect_robots, max_depth, suggest_feed.
- Dependencies: feedparser (robust RSS/Atom), defusedxml (XXE-safe XML).
- Docs: zh/en resources API, MCP/CLI/SDK help, ov.conf.example.
- Tests: 52 unit tests (fake httpx, no network).
2026-06-26 21:13:03 +08:00
Evo 27731e965f fix(parse): recognize .jsonl as a text file so upload encoding normalization applies (#2794)
Completes #2745, which added .jsonl to the vectorization text-extension set in
embedding_utils.py but left the parallel upload-time encoding path treating
.jsonl as non-text. is_text_file() decides text-vs-binary by exact suffix
membership across CODE_EXTENSIONS + DOCUMENTATION_EXTENSIONS +
ADDITIONAL_TEXT_EXTENSIONS, which had .json but not .jsonl (the suffix of
data.jsonl is .jsonl, not .json). So detect_and_convert_encoding skipped UTF-8
normalization for a legacy-encoded .jsonl -- unlike .json -- which then got
vectorized as text, the exact mojibake class #2770 fixed.

Add .jsonl to ADDITIONAL_TEXT_EXTENSIONS (next to .json); is_text_file unions
all three sets so one entry suffices. Behavior for every other extension is
unchanged. Adds a test assert. Refs #2745, #2744, #2770.
2026-06-25 15:46:54 +08:00
87329714dd feat(grep): integrate VikingDB bm25 keyword search for grep engine (#2144)
* feat(grep): integrate VikingDB bm25 keyword search for grep engine

* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)

* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison

* fix(schema): upsert data to vikingdb lack of content

* chore: add benchmark for retrieval

* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs

* fix(benchmark): sub uri args; add report

* refactor: code format by ruff

* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf

* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search

* fix: adjust benchmark scripts

* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls

* refactor: new benchmark

* fix: step1 add resource by real code data

* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex

* optimize (benchmark): adjust keywords and ground truth for testing

* fix: truncate 64KB for content field

* optimize: effectiveness add resource plainly

* optimize: change param use of SearchByKeywords from "keywords" to "query"

* optimize(benchmark): refactor effectiveness scripts

* optimize: ensure raw data for content field

* optimize: fulltext analyzer's stop-words only use symbols

* fix: adapt to new ov cli for benchmark

* optimize: reuse file content to avoid re-read AGFS file

* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts

* optimize: benchmark client timeout

* update README

* fix: rm unused param

* fix: default values in docs

* optimize: increase truncate byte size to 1MB for content field for VikingDB

* fix(logger): harden queued stream logging (#2786)

* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock

When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.

During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.

Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.

Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.

Closes: #2752

* fix(logger): harden queued stream logging

---------

Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
2026-06-24 18:46:02 +08:00
Hao Zhe 324f96ebb6 fix(parse): normalize legacy text encodings (#2770)
* fix(parse): normalize text file encodings

* fix(parser): harden text encoding normalization

* fix(parse): normalize text encodings with charset-normalizer

* test(parse): use synthetic gb18030 fixture text

* fix(parse): respect detector rank for non-cjk text

* fix(parse): rescue short simplified chinese text

* fix(parse): preserve korean hanja text

* style(parse): format text encoding tests
2026-06-23 16:55:32 +08:00
Evo 92646308e7 fix(parse): use safe_extract_zip for UnderstandingAPI zip download (#2634) 2026-06-17 21:05:11 +08:00