* feat(pdf): refactor MinerU parsing to the official file_parse API
* feat(pdf): remove mineru_api_key from configuration and examples
* feat(pdf): preflight MinerU /health during service initialization
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(embedding): downsample oversized image inputs
Keep imported image resources unchanged while avoiding provider-side multimodal embedding failures for oversized images. The embedding path now builds a temporary downsampled image data URI when image bytes exceed the shared large-image limits.
Move reusable image size thresholds into media_limits so both parser-side large image handling and embedding-side input preparation depend on a common utility instead of embedding_utils importing parser internals.
Add vectorize_file coverage confirming large image embedding inputs are resized and the stored resource bytes are preserved.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(media): downsample oversized image model inputs
Keep imported image resources unchanged while avoiding provider-side multimodal failures for oversized images. Shared image input preparation now builds temporary downsampled bytes for model requests when image bytes exceed the configured large-image limits.
Apply the model-input downsampling to both semantic image summary generation and embedding image data URI construction, so directory and code repository imports can preserve original images while sending provider-compatible inputs.
Move reusable image size thresholds into media_limits so parser-side large image handling, VLM summary generation, and embedding preparation share common limits without embedding_utils importing parser internals.
Always convert downsampled model images to RGB before JPEG encoding so Pillow-openable modes such as LA or I;16 do not fall back to the original oversized bytes.
Add coverage confirming VLM image summaries, vectorize_file embedding inputs, and JPEG-incompatible image modes are resized while stored resource bytes are preserved.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Frontmatter was parsed into ParseResult.meta and removed from the body, but
that metadata is never persisted, so every ingested markdown file lost its
frontmatter with no way to read the fields back.
Parse frontmatter into meta unconditionally (it still drives doc_title) and
only remove it from the stored body when explicitly configured; that removal
is now off by default.
_smart_split_content documents that it enforces both a token limit
(max_size) and a hard character limit. But when a single paragraph was
oversized by tokens yet under the character limit, the force-split loop
stepped through it by max_chars only, so it emitted a chunk that still
exceeded max_size tokens.
This is reachable with ordinary long-form CJK text: _estimate_token_count
weights CJK at ~0.7 token/char, so a ~5000-char Chinese paragraph is
~3500 tokens (over the 2048 default) while staying under the char limit,
and was returned as a single over-budget chunk.
Bound the force-split step by min(max_chars, max_size / MAX_TOKENS_PER_CHAR),
where MAX_TOKENS_PER_CHAR is the worst-case (CJK) density already used by
_estimate_token_count, now extracted into a shared constant so the two stay
in sync.
Add regression tests for the token budget and content preservation.
_gh_slug claims to produce "GitHub-style" heading slugs but collapsed runs
of whitespace (`re.sub(r"\s+", "-", s)`). GitHub's reference slugger
(github-slugger) maps each space to its own hyphen (`.replace(/ /g, '-')`)
and does not collapse.
Because punctuation is stripped before spaces are converted, a heading like
"Foo & Bar" leaves two adjacent spaces where "&" was. GitHub renders this as
"foo--bar", but _gh_slug produced "foo-bar". The intra-document link rewriter
(_rewrite_link) compares _gh_slug(heading) against the link fragment, so an
author-written link such as `guide.md#foo--bar` failed to match its heading
and was left unrewritten after the target doc was split into sections.
Replace `\s+` with `\s` so each whitespace character maps to one hyphen,
matching GitHub. Simple single-space headings are unaffected.
Add regression tests covering punctuation headings and ordinary headings.
* feat: add audio and video understanding via VLM
* docs: design media resource guards
* fix: bound media staging concurrency
* fix: cap unknown-size media staging
* test: stage media in routing fake
* test: exercise media staging callbacks
* test: trim media understanding coverage
* chore: 清理实现计划文档
* fix: 修复多凭证切换问题
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Print-to-PDF producers routinely emit several image XObjects drawn at the
exact same position on a page (a background layer plus a content layer).
Because `_extract_image_from_page` rasterises the page *region* rather than
decoding the XObject itself, every one of them renders to identical bytes —
so a document with two stacked full-page layers wrote two byte-identical
PNGs per page and referenced both from the generated markdown.
Dedup within each page, in two steps:
- bbox first, so a repeat is skipped before paying for the render;
- a content hash as a backstop, for bboxes that differ slightly but still
rasterise to the same bytes.
Both sets are per-page, so a header logo repeated across pages is still
kept once on every page. `meta["images_deduplicated"]` reports how many
were skipped.
Measured on an 8-page article exported from a web page: 16 saved PNGs -> 8,
16 markdown image references -> 8, local conversion 3.8s -> 2.4s.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(parse): distinguish mpegts from TypeScript ts
* fix(parse): tighten mpegts ts routing semantics
* fix(semantic): use file name for media summary type
---------
Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* refactor(parse): simplify resource ingestion routing
Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.
* fix(feishu): preserve sheet and bitable imports
Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.
* fix(parse): keep normalized Feishu content internal
Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.
* fix(feishu): parse bitable blocks embedded in sheets
Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.
* fix(feishu): download bitable attachment images
* refactor(parse): remove unused document converter
* refactor(parse): unify Understanding routing
* docs(parse): mark routing classification points
* docs(parse): complete wait routing flow
* fix(parse): preserve Feishu Base URL scope
* refactor(resource): separate ingestion submission from execution
* fix(resource): reject internal ingestion fields at public entry
* fix(parse): handle parentheses in Markdown image paths
The image regex !\[([^\]]*)\]\(([^)]+)\) used [^)]+ for the path capture
group, which truncates at the first ) character. When document titles
or filenames contain balanced parentheses (e.g. "文档_17 (17号项目)"), the
generated image paths include ) and the regex captures a truncated,
non-existent path. This causes _resolve_image_path() to fail silently
(WARNING only), and the image is never copied to VikingFS or sent to
VLM for understanding.
Fix: replace the path capture group with (?:[^()]|\([^()]*\))+, which
allows one level of balanced parentheses inside the path while still
terminating at the correct closing ) of the Markdown image syntax.
Add focused tests covering balanced parens in directory and filename
components, URLs with parens, multiple images on one line, and
non-matching of plain links.
Fixes#3455
* fix(test): exercise MarkdownParser._image_pattern directly, remove unused import
Address review feedback on #3462:
1. Tests now import and instantiate MarkdownParser to access the
production _image_pattern regex, instead of compiling an independent
copy. Tests fail if the production regex regresses.
2. Remove unused `import pytest` (Ruff F401).
* fix(parse): rewrite parenthesized image paths
* test: remove extra image rewrite regression case
* fix(parse): share markdown image parsing for rewrite
* refactor(parse): keep markdown image fix minimal
---------
Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(parse): add large image processing for image parser
- Add large_image_processor.py: detect large images (>10MB or >4096px),
create low-res previews, split into grid tiles, and generate grid
overlay images with tile labels
- Refactor ImageParser.parse() to integrate large image processing pipeline
- Enable SVG-to-PNG conversion in utils.py (cairosvg/wand)
- Rename ImageConfig.max_dimension to preview_max_dimension and add new
config fields: max_file_size_mb, max_tile_size_mb, max_tile_dimension_px,
tile_overlap_px, large_image_threshold_dimension
- Update ov.conf.example with new image config options
* fix(parse): correct tile dimension comment from 1024px to 2048px
* fix(parse): fix tile label path in grid overlay to include tiles/ directory
* fix(parse): register missing image extensions for ImageParser
TIFF, ICO, DIB, ICNS, SGI, JP2 were not in IMAGE_EXTENSIONS, causing
them to fallback to TextParser. All are supported by PIL.
* fix(parse): preserve PNG format for tiles instead of always converting to JPEG
* fix(parse): address review feedback for large image processing
- Wire config.image to ImageParser in ParserRegistry (was missing)
- Remove unnecessary preview creation for small images (broke LA mode PNG)
- Enforce max_tile_size_mb on tiles with quality reduction and resize fallback
- Remove 64-tile hard cap that conflicted with max_tile_dimension_px
- Add comment explaining why original file is not saved for large images
* refactor(parse): remove max_tile_size_mb as it is a soft suggestion
max_tile_size_mb was a soft constraint that was not enforced
consistently. Remove it from config, constants, and all enforcement
logic. Tile dimension (max_tile_dimension_px) remains the sole constraint.
* fix(parse): use CJK-capable font for grid overlay labels
The old font loading only tried macOS-specific paths and fell back to
PIL's default bitmap font, which cannot render CJK characters in
filenames. Add a cross-platform CJK font lookup that covers Linux
(Noto/Droid/WQY/DejaVu), macOS (PingFang), and Windows (MSYH/SimSun).
* fix(parse): convert non-VLM-supported image formats to PNG on save
Image formats like TIFF, ICO, DIB, ICNS, SGI, JP2 are not recognized
by VLM backends (OpenAI/LiteLLM/VolcEngine only support PNG/JPEG/GIF/
WebP/BMP) or by embedding_utils for image vectorization. When a file
with one of these extensions is parsed, convert it to PNG and use a
.png extension so that downstream pipelines see consistent data.
SVG files (already PNG-converted via cairosvg) also get the .png
extension for the same reason.
* fix(parse): import io for SVG conversion
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
中文:对 HTTP/HTTPS URL 使用 urlparse(url).path 提取扩展名,确保 video.mp4?signature=... 命中 UnderstandingAPI fast path。新增带签名视频 URL 的 ParserRouter 回归测试。
English: Parse HTTP/HTTPS URL paths before checking extensions so signed video URLs hit the UnderstandingAPI fast path. Add a ParserRouter regression test for signed video URLs.
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
When the server crawls a URL that returns 401/403, it wraps the failure
in a 5xx envelope. The CLI matched on auth-flavored message text alone
and rendered "OpenViking rejected the API key", wrongly pointing users
at their local config.
Gate the API-key error report on status: 5xx responses are never treated
as a client auth failure even when the message mentions authentication or
forbidden. Also split the 401 vs 403 fetch messages so 403 reads as an
access-denied / anti-bot block rather than a credential problem.
* fix(feishu): surface permission errors clearly and keep users on page
Map Feishu/Lark API failures to typed OpenViking errors with actionable hints, and keep Web Studio from treating HTTP 403 permission denials as session logout.
* fix(feishu): simplify API error mapping
* refactor(feishu): inline API error mapping
---------
Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(feishu): persist imported document images
Download Feishu image tokens into local import temp trees so Markdown image references can be ingested alongside the document.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(feishu): harden inline-image download (async, extension, user token)
Address review feedback on the inline-image import path:
- Run the synchronous lark-oapi media download via asyncio.to_thread so a
slow Feishu request no longer blocks unrelated async work on the event loop,
matching the existing _fetch_document() pattern.
- Infer the image file extension from the downloaded bytes (byte-magic
sniffing) and fall back to the response Content-Type via the existing
mime_types.get_preferred_extension helper, instead of hardcoding .png. This
stops JPEG/WebP/GIF bytes from being mislabeled as PNG to downstream
consumers (e.g. the data:image/... URI built during multimodal vectorization).
- Advertise AccessTokenType.USER on the media download request when a user
access token is supplied, so lark-oapi actually injects it. Previously the
request only allowed TENANT, so user-token imports read the document body
but silently dropped every image.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(storage): optimize glob func
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* fix(localfs): offload blocking fs operations to spawn_blocking
* feat(glob): cap glob api default node_limit at 256
* feat(sdk): add node_limit options for glob in python and go SDKs
* refactor(web-crawler): remove Playwright rendering, keep static SPA shells
Drop the Playwright-based fallback rendering path so SPA pages are stored
as their static shell HTML instead of being rendered headlessly. SPA
shells now surface the <noscript> notice (e.g. "You need to enable
JavaScript to run this app.") rather than producing an empty document,
and robots.txt-blocked entry pages get a human-readable error message.
- Delete playwright_renderer.py, render_heuristics.py and their tests
- Strip fallback_playwright/playwright_timeout config, fallback_rendered
counter, and render-hint plumbing from the crawler and web importer
- HTMLParser: drop the SPA-empty-pattern stripping and fall back to
<noscript> text when trafilatura extracts nothing
- Humanize the robots.txt entry-failure message
- Migrate the orphaned _convert_to_raw_url tests to HTTPAccessor (the
method moved there in an earlier reorg) and remove the stale file
* refactor(web-importer): simplify robots.txt-blocked import message
Replace the verbose robots.txt explanation with a short compliance hint
that omits the URL and points users to local-file import instead.
* Refactor recursive web import into HTTP accessor
Move ordinary web page import routing into HTTPAccessor and materialize crawled pages as a temporary directory via WebImporter.
Relocate Scrapy/Playwright crawling under parse.accessors.web_crawler, keep trafilatura extraction inside HTMLParser, and avoid repeated ResourceService.add_resource calls.
Add recursive crawl controls, safe request validation, page/download classification, and focused unit coverage.
* Document recursive web crawler options
* fix(web-crawler): stop SSRF sub-resource block from failing whole render
The playwright fallback validated every sub-resource request against the
SSRF guard and raised on the first disallowed host, failing the entire
page render. volcengine docs load a probe resource on an internal host,
so rendering always failed and the crawler stored the static anti-bot
"Please wait..." challenge page as content.
Now a blocked sub-resource is only aborted; the main document and final
URL still gate the result. Also wait past JS interstitials, retry reads
through in-flight navigation, and reject shell/challenge pages instead of
storing them.
* fix(web-crawler): surface renderer error hint on entry-page failure
When Playwright is unavailable, the renderer returns an actionable install
hint via RenderResult.error, but the spider silently kept the static shell
and WebImporter raised only the generic "Failed to fetch entry page". The
hint never reached the user.
Now the spider records rendered.error on the failed page, and WebImporter
appends the entry page's failure reason to the raised message so the CLI
shows the Playwright install instructions.
* fix(web-crawler): surface render hints and enforce crawl limits
* fix(web-crawler): avoid rendering SSR app pages
* perf(web-crawler): bound render concurrency and cap networkidle wait
Playwright renders were dispatched from parse callbacks without any
concurrency limit, so a page with many child links could spawn dozens of
Chromium pages at once (observed peak 28 for a 20-page crawl), risking OOM
on large sites and starting ~2.3x more renders than needed before
max_pages stopped the crawl. Gate renders with a semaphore sized to
config.concurrency and re-check the success limit after acquiring a slot
so queued callbacks skip rendering once the crawl is already done.
Also cap the networkidle wait at 8s: pages with continuous background
activity (e.g. GraphiQL) never go idle and previously blocked until the
full render timeout, turning a ~3s page into ~38s. Content is ready after
domcontentloaded and _wait_past_challenge covers late-arriving text.
Bump default concurrency 5 -> 10.
* fix(web-crawler): route .html/.htm URLs through recursive WebImporter
An explicit .html/.htm URL is detected as DOWNLOAD_HTML via the extension
map, so access() previously only routed URLType.WEBPAGE to WebImporter and
these URLs fell through to single-file download, silently ignoring
depth/max_pages. Route DOWNLOAD_HTML through WebImporter too, treating a
single-page import as the depth=0 case.
* fix(web-crawler): improve HTML extraction and rendering heuristics
- Drop trafilatura favor_precision=True: it stripped the full body of
link-dense pages, keeping only headers.
- Only render __NEXT_DATA__ pages with Playwright when their static body
is too thin; SSR/SSG Next.js pages already ship full text.
- Disable Scrapy telnet console to avoid opening port 6023.
* fix(web-crawler): keep code-hosting single-file URLs off recursive crawler
GitHub/GitLab blob and GitHub raw URLs resolve to a single file, not a
site. Route them through the single-file download path instead of the
recursive WebImporter, which otherwise crawls the hosting UI shell.
* docs(resources): add recursive web crawler usage examples
Add depth/max_pages crawl examples to the HTTP, Python SDK, and CLI
blocks in both the zh and en resource API docs, plus path-prefix
filtering and skip_download_links variants.
Add WebFeedAccessor (priority 60) that turns a single sitemap /
sitemapindex / RSS / Atom URL into ONE resource tree: it mirrors every
listed page into a temp directory and reuses the existing DirectoryParser
pipeline (the same "fetch-many -> dir -> tree" contract as GitAccessor).
A watch on the feed URL keeps the whole site refreshed (new pages added,
removed pages dropped on each rebuild).
- New openviking/parse/accessors/web_feed_accessor.py: WebFeedAccessor +
sitemap/feed extractors (nested sitemapindex recursion with depth cap,
RSS 2.0 / Atom via feedparser), bounded concurrent polite mirroring,
robots.txt, same-host / include / exclude / max_pages limits.
- args={"site": true} forces whole-site ingestion from a bare domain or
page by auto-discovering the sitemap/RSS (robots.txt, HTML
<link rel=alternate>, conventional paths); {"site": false} opts a
feed-looking URL back out to HTTPAccessor.
- Thread accessor-selection kwargs through can_handle; the registry
tolerates accessors whose can_handle lacks **kwargs (back-compatible).
- Single-page adds get a non-blocking "this site exposes a sitemap/RSS"
suggestion appended to the MCP add_resource response, gated to the
site root only; never auto-crawls.
- New WebFeedConfig (parsers.webfeed): max_pages, concurrency, politeness
delay, same_host_only, respect_robots, max_depth, suggest_feed.
- Dependencies: feedparser (robust RSS/Atom), defusedxml (XXE-safe XML).
- Docs: zh/en resources API, MCP/CLI/SDK help, ov.conf.example.
- Tests: 52 unit tests (fake httpx, no network).
Completes #2745, which added .jsonl to the vectorization text-extension set in
embedding_utils.py but left the parallel upload-time encoding path treating
.jsonl as non-text. is_text_file() decides text-vs-binary by exact suffix
membership across CODE_EXTENSIONS + DOCUMENTATION_EXTENSIONS +
ADDITIONAL_TEXT_EXTENSIONS, which had .json but not .jsonl (the suffix of
data.jsonl is .jsonl, not .json). So detect_and_convert_encoding skipped UTF-8
normalization for a legacy-encoded .jsonl -- unlike .json -- which then got
vectorized as text, the exact mojibake class #2770 fixed.
Add .jsonl to ADDITIONAL_TEXT_EXTENSIONS (next to .json); is_text_file unions
all three sets so one entry suffices. Behavior for every other extension is
unchanged. Adds a test assert. Refs #2745, #2744, #2770.
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* fix(parse): normalize text file encodings
* fix(parser): harden text encoding normalization
* fix(parse): normalize text encodings with charset-normalizer
* test(parse): use synthetic gb18030 fixture text
* fix(parse): respect detector rank for non-cjk text
* fix(parse): rescue short simplified chinese text
* fix(parse): preserve korean hanja text
* style(parse): format text encoding tests