Commit Graph
88 Commits
Author SHA1 Message Date
baojun-zhang 84c0895c44 fix(queuefs): skip add-resource lock replay after persisted result (#4007) 2026-08-14 17:12:38 +08:00
MaojiaShengandTRAE CLI 5aed7f72b4 fix(media): downsample oversized image model inputs (#3965)
* fix(embedding): downsample oversized image inputs

Keep imported image resources unchanged while avoiding provider-side multimodal embedding failures for oversized images. The embedding path now builds a temporary downsampled image data URI when image bytes exceed the shared large-image limits.

Move reusable image size thresholds into media_limits so both parser-side large image handling and embedding-side input preparation depend on a common utility instead of embedding_utils importing parser internals.

Add vectorize_file coverage confirming large image embedding inputs are resized and the stored resource bytes are preserved.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(media): downsample oversized image model inputs

Keep imported image resources unchanged while avoiding provider-side multimodal failures for oversized images. Shared image input preparation now builds temporary downsampled bytes for model requests when image bytes exceed the configured large-image limits.

Apply the model-input downsampling to both semantic image summary generation and embedding image data URI construction, so directory and code repository imports can preserve original images while sending provider-compatible inputs.

Move reusable image size thresholds into media_limits so parser-side large image handling, VLM summary generation, and embedding preparation share common limits without embedding_utils importing parser internals.

Always convert downsampled model images to RGB before JPEG encoding so Pillow-openable modes such as LA or I;16 do not fall back to the original oversized bytes.

Add coverage confirming VLM image summaries, vectorize_file embedding inputs, and JPEG-incompatible image modes are resized while stored resource bytes are preserved.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-12 22:24:57 +08:00
c1345a1f7e feat(feishu):Support Feishu Drive folder and file imports (#3937)
* Support Lark drive folder and file URLs

* Handle partial Feishu folder import failures

* Fix Feishu drive folder path names

* Tighten Feishu drive folder tests

* test(feishu): consolidate drive import coverage

---------

Co-authored-by: haoxingjun <haoxingjun@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-12 16:01:00 +08:00
Zayn Jarvis dcca29364c fix(parse): stop silently dropping markdown YAML frontmatter (#3929)
Frontmatter was parsed into ParseResult.meta and removed from the body, but
that metadata is never persisted, so every ingested markdown file lost its
frontmatter with no way to read the fields back.

Parse frontmatter into meta unconditionally (it still drives doc_title) and
only remove it from the stored body when explicitly configured; that removal
is now off by default.
2026-08-11 14:14:53 +08:00
bianbiandashen 3087f943a2 fix(markdown): keep force-split chunks within the token budget (#3672)
_smart_split_content documents that it enforces both a token limit
(max_size) and a hard character limit. But when a single paragraph was
oversized by tokens yet under the character limit, the force-split loop
stepped through it by max_chars only, so it emitted a chunk that still
exceeded max_size tokens.

This is reachable with ordinary long-form CJK text: _estimate_token_count
weights CJK at ~0.7 token/char, so a ~5000-char Chinese paragraph is
~3500 tokens (over the 2048 default) while staying under the char limit,
and was returned as a single over-budget chunk.

Bound the force-split step by min(max_chars, max_size / MAX_TOKENS_PER_CHAR),
where MAX_TOKENS_PER_CHAR is the worst-case (CJK) density already used by
_estimate_token_count, now extracted into a shared constant so the two stay
in sync.

Add regression tests for the token budget and content preservation.
2026-08-08 01:54:06 +08:00
bianbiandashen 7b8b33e8f8 fix(markdown): match GitHub anchor slugs for headings with punctuation (#3673)
_gh_slug claims to produce "GitHub-style" heading slugs but collapsed runs
of whitespace (`re.sub(r"\s+", "-", s)`). GitHub's reference slugger
(github-slugger) maps each space to its own hyphen (`.replace(/ /g, '-')`)
and does not collapse.

Because punctuation is stripped before spaces are converted, a heading like
"Foo & Bar" leaves two adjacent spaces where "&" was. GitHub renders this as
"foo--bar", but _gh_slug produced "foo-bar". The intra-document link rewriter
(_rewrite_link) compares _gh_slug(heading) against the link fragment, so an
author-written link such as `guide.md#foo--bar` failed to match its heading
and was left unrewritten after the target doc was split into sections.

Replace `\s+` with `\s` so each whitespace character maps to one hyphen,
matching GitHub. Simple single-space headings are unaffected.

Add regression tests covering punctuation headings and ordinary headings.
2026-08-07 20:56:40 +08:00
Kchen 8d1d52fe5d 资源导入:支持解析后不拆分文档 (#3645) 2026-08-05 11:34:10 +08:00
Haoyu ZhangandQin Haojie 3f3554256b feat: 支持基于火山方舟的音视频多模态理解 (#3563)
* feat: add audio and video understanding via VLM

* docs: design media resource guards

* fix: bound media staging concurrency

* fix: cap unknown-size media staging

* test: stage media in routing fake

* test: exercise media staging callbacks

* test: trim media understanding coverage

* chore: 清理实现计划文档

* fix: 修复多凭证切换问题

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-08-03 16:44:22 +08:00
yangxinxin-7andClaude Opus 5 8b4deaab99 fix(parser): drop duplicate images when a PDF stacks XObjects on one spot (#3662)
Print-to-PDF producers routinely emit several image XObjects drawn at the
exact same position on a page (a background layer plus a content layer).
Because `_extract_image_from_page` rasterises the page *region* rather than
decoding the XObject itself, every one of them renders to identical bytes —
so a document with two stacked full-page layers wrote two byte-identical
PNGs per page and referenced both from the generated markdown.

Dedup within each page, in two steps:

- bbox first, so a repeat is skipped before paying for the render;
- a content hash as a backstop, for bboxes that differ slightly but still
  rasterise to the same bytes.

Both sets are per-page, so a header logo repeated across pages is still
kept once on every page. `meta["images_deduplicated"]` reports how many
were skipped.

Measured on an 8-page article exported from a web page: 16 saved PNGs -> 8,
16 markdown image references -> 8, local conversion 3.8s -> 2.4s.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:03:21 +08:00
Qin Haojie 09e42aa739 fix(resource): restore queue status for waited imports (#3658) 2026-07-31 15:40:27 +08:00
zgy 49b182045b refactor(parser): Refactor code summaries to fixed skeleton-first routing (#3568)
* Refactor code summary skeleton routing

* Simplify code skeleton routing configuration

* Render C tag skeletons as signatures

* Revert "Render C tag skeletons as signatures"

This reverts commit 8e342055f8.

* Simplify fixed code skeleton summary route

* Inline process skeleton rendering

* Simplify code skeleton routing entrypoints

* Fix code summary review issues

* Address final code summary review feedback

* Route failed tags skeletons to LLM fallback

* Restore CUDA and TS extension routing

* Improve code skeleton query coverage

* Route semantic code detection through skeleton support

* Move process skeleton engine into ast package

* Admit skeleton-supported files during directory scan

* Align code summary docs after main merge

* Reduce code skeleton fallback log verbosity

* chore: require grep-ast 0.9.0
2026-07-31 11:38:57 +08:00
Qin Haojie fd42b1ad92 feat(tasks): support task cancellation (#3577)
* feat(tasks): support task cancellation

* refactor(tasks): scope cancellation to current user

* feat(cli): support task cancellation

* refactor(tasks): make cancellation queue-aware

* refactor(tasks): simplify cancellation bookkeeping

* test: remove task cancellation coverage

* refactor(tasks): trim cancellation coordination

* fix(tasks): contain cancellation to owned work

* feat(tasks): persist resource source metadata

* fix(tasks): handle cancelled work consistently

* refactor(tasks): make completion queue-aware

* fix(tasks): persist terminal state before queue ack

* test(tasks): remove added lifecycle tests

* docs(tasks): document task cancellation
2026-07-30 20:34:27 +08:00
Eurakaxun 44c6df2622 perf: retrieval, import, LangChain, and session-context optimizations (#3569) 2026-07-30 10:10:11 +08:00
baojun-zhang 2f9451231e refactor(pathlock):using rust implement instead python (#3602)
* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):optimize unit test code

* refactor(pathlock):optimize encryption create func

* refactor(pathlock):avoid releasing handoffed pathlock on enqueue errors

* fix(pathlock): use owned lease capability and handle S3 create-new 409 as conflict

* fix(ragfs): keep original FsContext for multi-write metadata

* fix(pathlock): resolve lease coverage and CAS handling issues

- detect S3 conditional conflicts from structured service errors
- pass transaction leases when deleting skill roots
- let temp cleanup acquire locks for temp paths
- disambiguate cache and pathlock providers in cache tests
- update temp cleanup lease assertions

* fix(ragfs): bypass pathlock for multi-write metadata

* fix(ragfs): revert pathlock fail-fast design

* fix(ragfs):fix(ragfs): use non-blocking fcntl locks for localfs CAS

* fix(ragfs): serialize heartbeat lease refresh with release and report real conflict kind

* fix(ragfs): preserve conflict kind snapshot and drop unused test scaffolding

* fix(ragfs): preserve conflict kind snapshot and drop unused test scaffolding
2026-07-29 19:45:34 +08:00
zihengli e9c4cc97c3 refactor: extract Connector delegation and expose declarative add_type (#3591)
* refactor: delegate add_resource imports to external Connector

* refactor: delegate add_resource imports to external Connector

* refactor: extract Connector delegation and expose declarative add_type

* fix: merge main to refactor/connector_delegator

* fix: merge main to refactor/connector_delegator
2026-07-29 18:09:25 +08:00
chenxiaobin-monkeyandchenxiaobin.monkey ff37e25cfd fix(parse): distinguish mpegts from TypeScript ts (#3574)
* fix(parse): distinguish mpegts from TypeScript ts

* fix(parse): tighten mpegts ts routing semantics

* fix(semantic): use file name for media summary type

---------

Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
2026-07-29 13:28:25 +08:00
baojun-zhang 1841dfed81 Revert "refactor(pathlock):using rust implement instead python (#3557)" (#3597)
This reverts commit 6b538db569.
2026-07-29 11:31:41 +08:00
baojun-zhang 6b538db569 refactor(pathlock):using rust implement instead python (#3557)
* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):optimize unit test code

* refactor(pathlock):optimize encryption create func

* refactor(pathlock):avoid releasing handoffed pathlock on enqueue errors

* fix(pathlock): use owned lease capability and handle S3 create-new 409 as conflict

* fix(ragfs): keep original FsContext for multi-write metadata

* fix(pathlock): resolve lease coverage and CAS handling issues

- detect S3 conditional conflicts from structured service errors
- pass transaction leases when deleting skill roots
- let temp cleanup acquire locks for temp paths
- disambiguate cache and pathlock providers in cache tests
- update temp cleanup lease assertions

* fix(ragfs): bypass pathlock for multi-write metadata

* fix(ragfs): revert pathlock fail-fast design
2026-07-29 11:08:42 +08:00
Jiahui Zhou 5d1ba45be4 Feat/add resource processing mode (#3566)
* feat: add resource processing mode

* fix: keep semantic artifacts in vectors-only add resource

* test: support processing mode in api test client

* docs: document add resource processing mode

* fix: align processing mode after resource ingestion refactor

* feat: expose processing mode in TypeScript SDK

* fix: preserve add resource compatibility
2026-07-28 20:09:06 +08:00
Evo 16e7c33f1f fix(feishu): preserve Bitable media permission context during import (#3558) 2026-07-28 11:41:39 +08:00
Qin Haojie 2be4bb4879 refactor(parse): 收口资源解析路由 (#3295)
* refactor(parse): simplify resource ingestion routing

Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.

* fix(feishu): preserve sheet and bitable imports

Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.

* fix(parse): keep normalized Feishu content internal

Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.

* fix(feishu): parse bitable blocks embedded in sheets

Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.

* fix(feishu): download bitable attachment images

* refactor(parse): remove unused document converter

* refactor(parse): unify Understanding routing

* docs(parse): mark routing classification points

* docs(parse): complete wait routing flow

* fix(parse): preserve Feishu Base URL scope

* refactor(resource): separate ingestion submission from execution

* fix(resource): reject internal ingestion fields at public entry
2026-07-27 16:47:20 +08:00
Wu JiaCheng 5ab30d3e04 fix(parse): preserve text file parser metadata (#3480) 2026-07-24 16:30:38 +08:00
Qin Haojie fd098cfd65 fix(cli): preserve structured API errors (#3379)
* fix(error): preserve structured errors in cli

* fix(cli): preserve status for non-json errors
2026-07-22 18:06:48 +08:00
Wu JiaCheng 061359a2f5 fix(parse): normalize MIME aliases with parameters (#3393) 2026-07-22 15:44:43 +08:00
Wu JiaCheng 63c878e60e fix(parse): import extensionless README files (#3394) 2026-07-22 15:29:13 +08:00
Qin Haojie 339817d981 fix(resource): isolate external parsing queue (#3448)
Keep local resource ingestion from waiting behind UnderstandingAPI jobs while preserving durable queue recovery.
2026-07-22 15:11:50 +08:00
ShaoZegangByte 39dc01e2a5 feat(resource): support Feishu/Lark URL imports via UnderstandingAPI (#3320)
* feat(resource):add lark understand api

* fix(resource):add lark understand api env

* fix(resource):add lark understand api

* fix(resource):understand api pr

* fix(resource):add test

* fix(resource):handle deferred URIs for async Feishu imports

* fix(resource):fix Ruff issues in lark import

* fix: persist final resource URI for async UnderstandingAPI tasks

* fix: clean up cancelled async UnderstandingAPI scheduling

* fix: getattr defer_target_resolution
2026-07-21 11:08:13 +08:00
huangruitengandhuangruiteng f093fafd1b fix(feishu): preserve title prefixes in resource names (#3366)
Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-20 17:16:49 +08:00
Qin Haojie d47f2106ee refactor: remove unused and deprecated APIs (#3272)
Delete dead compatibility paths and test-only helpers so unsupported APIs do not remain as accidental contracts.
2026-07-16 10:49:56 +08:00
Haoyu Zhangandzhanghaoyu c6d48bc056 feat(parse): support AC-3 audio resources and add multimodal integration tests (#3229)
* docs: design AC-3 resource ingestion

* chore: ignore local worktrees

* feat(parse): support AC-3 audio resources

* feat: 补充多模态文档解析测试脚本

---------

Co-authored-by: zhanghaoyu <zhanghaoyu.la@bytedance.com>
2026-07-14 11:03:49 +08:00
huangruitengandhuangruiteng cbcec52d7d fix(parse): preserve filters for local git repositories (#3190)
Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-13 11:35:21 +08:00
Hao Zhe 47170b05dc fix(parse): route mislabeled OOXML Word files correctly (#3113) 2026-07-10 14:29:59 +08:00
Kchenandchenpengfei e19d5f7d66 修复签名视频链接的解析路由 / Fix signed video URL parser routing (#3103)
中文:对 HTTP/HTTPS URL 使用 urlparse(url).path 提取扩展名,确保 video.mp4?signature=... 命中 UnderstandingAPI fast path。新增带签名视频 URL 的 ParserRouter 回归测试。

English: Parse HTTP/HTTPS URL paths before checking extensions so signed video URLs hit the UnderstandingAPI fast path. Add a ParserRouter regression test for signed video URLs.

Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
2026-07-09 20:34:20 +08:00
cd9add7a27 fix(feishu): surface permission errors clearly and keep users on page (#3032)
* fix(feishu): surface permission errors clearly and keep users on page

Map Feishu/Lark API failures to typed OpenViking errors with actionable hints, and keep Web Studio from treating HTTP 403 permission denials as session logout.

* fix(feishu): simplify API error mapping

* refactor(feishu): inline API error mapping

---------

Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-07 19:55:43 +08:00
07aa9dc775 feat(feishu): persist imported document images (#3033)
* feat(feishu): persist imported document images

Download Feishu image tokens into local import temp trees so Markdown image references can be ingested alongside the document.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(feishu): harden inline-image download (async, extension, user token)

Address review feedback on the inline-image import path:

- Run the synchronous lark-oapi media download via asyncio.to_thread so a
  slow Feishu request no longer blocks unrelated async work on the event loop,
  matching the existing _fetch_document() pattern.
- Infer the image file extension from the downloaded bytes (byte-magic
  sniffing) and fall back to the response Content-Type via the existing
  mime_types.get_preferred_extension helper, instead of hardcoding .png. This
  stops JPEG/WebP/GIF bytes from being mislabeled as PNG to downstream
  consumers (e.g. the data:image/... URI built during multimodal vectorization).
- Advertise AccessTokenType.USER on the media download request when a user
  access token is supplied, so lark-oapi actually injects it. Previously the
  request only allowed TENANT, so user-token imports read the document body
  but silently dropped every image.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 17:44:26 +08:00
baojun-zhang 6a33ebb7ca Optimize glob walkdir (#3013)
* feat(storage): optimize glob func

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* fix(localfs): offload blocking fs operations to spawn_blocking

* feat(glob): cap glob api default node_limit at 256

* feat(sdk): add node_limit options for glob in python and go SDKs
2026-07-06 21:39:16 +08:00
zgy 8d861fabfe fix(web-crawler): remove Playwright rendering andoptimize the robots.txt entry-failure message (#3040)
* refactor(web-crawler): remove Playwright rendering, keep static SPA shells

Drop the Playwright-based fallback rendering path so SPA pages are stored
as their static shell HTML instead of being rendered headlessly. SPA
shells now surface the <noscript> notice (e.g. "You need to enable
JavaScript to run this app.") rather than producing an empty document,
and robots.txt-blocked entry pages get a human-readable error message.

- Delete playwright_renderer.py, render_heuristics.py and their tests
- Strip fallback_playwright/playwright_timeout config, fallback_rendered
  counter, and render-hint plumbing from the crawler and web importer
- HTMLParser: drop the SPA-empty-pattern stripping and fall back to
  <noscript> text when trafilatura extracts nothing
- Humanize the robots.txt entry-failure message
- Migrate the orphaned _convert_to_raw_url tests to HTTPAccessor (the
  method moved there in an earlier reorg) and remove the stale file

* refactor(web-importer): simplify robots.txt-blocked import message

Replace the verbose robots.txt explanation with a short compliance hint
that omits the URL and points users to local-file import instead.
2026-07-06 18:48:39 +08:00
zgy a50e9fd677 feat: add recursive web crawler based on Scrapy (#2836)
* Refactor recursive web import into HTTP accessor

Move ordinary web page import routing into HTTPAccessor and materialize crawled pages as a temporary directory via WebImporter.

Relocate Scrapy/Playwright crawling under parse.accessors.web_crawler, keep trafilatura extraction inside HTMLParser, and avoid repeated ResourceService.add_resource calls.

Add recursive crawl controls, safe request validation, page/download classification, and focused unit coverage.

* Document recursive web crawler options

* fix(web-crawler): stop SSRF sub-resource block from failing whole render

The playwright fallback validated every sub-resource request against the
SSRF guard and raised on the first disallowed host, failing the entire
page render. volcengine docs load a probe resource on an internal host,
so rendering always failed and the crawler stored the static anti-bot
"Please wait..." challenge page as content.

Now a blocked sub-resource is only aborted; the main document and final
URL still gate the result. Also wait past JS interstitials, retry reads
through in-flight navigation, and reject shell/challenge pages instead of
storing them.

* fix(web-crawler): surface renderer error hint on entry-page failure

When Playwright is unavailable, the renderer returns an actionable install
hint via RenderResult.error, but the spider silently kept the static shell
and WebImporter raised only the generic "Failed to fetch entry page". The
hint never reached the user.

Now the spider records rendered.error on the failed page, and WebImporter
appends the entry page's failure reason to the raised message so the CLI
shows the Playwright install instructions.

* fix(web-crawler): surface render hints and enforce crawl limits

* fix(web-crawler): avoid rendering SSR app pages

* perf(web-crawler): bound render concurrency and cap networkidle wait

Playwright renders were dispatched from parse callbacks without any
concurrency limit, so a page with many child links could spawn dozens of
Chromium pages at once (observed peak 28 for a 20-page crawl), risking OOM
on large sites and starting ~2.3x more renders than needed before
max_pages stopped the crawl. Gate renders with a semaphore sized to
config.concurrency and re-check the success limit after acquiring a slot
so queued callbacks skip rendering once the crawl is already done.

Also cap the networkidle wait at 8s: pages with continuous background
activity (e.g. GraphiQL) never go idle and previously blocked until the
full render timeout, turning a ~3s page into ~38s. Content is ready after
domcontentloaded and _wait_past_challenge covers late-arriving text.

Bump default concurrency 5 -> 10.

* fix(web-crawler): route .html/.htm URLs through recursive WebImporter

An explicit .html/.htm URL is detected as DOWNLOAD_HTML via the extension
map, so access() previously only routed URLType.WEBPAGE to WebImporter and
these URLs fell through to single-file download, silently ignoring
depth/max_pages. Route DOWNLOAD_HTML through WebImporter too, treating a
single-page import as the depth=0 case.

* fix(web-crawler): improve HTML extraction and rendering heuristics

- Drop trafilatura favor_precision=True: it stripped the full body of
  link-dense pages, keeping only headers.
- Only render __NEXT_DATA__ pages with Playwright when their static body
  is too thin; SSR/SSG Next.js pages already ship full text.
- Disable Scrapy telnet console to avoid opening port 6023.

* fix(web-crawler): keep code-hosting single-file URLs off recursive crawler

GitHub/GitLab blob and GitHub raw URLs resolve to a single file, not a
site. Route them through the single-file download path instead of the
recursive WebImporter, which otherwise crawls the hosting UI shell.

* docs(resources): add recursive web crawler usage examples

Add depth/max_pages crawl examples to the HTTP, Python SDK, and CLI
blocks in both the zh and en resource API docs, plus path-prefix
filtering and skip_download_links variants.
2026-07-03 19:25:10 +08:00
87329714dd feat(grep): integrate VikingDB bm25 keyword search for grep engine (#2144)
* feat(grep): integrate VikingDB bm25 keyword search for grep engine

* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)

* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison

* fix(schema): upsert data to vikingdb lack of content

* chore: add benchmark for retrieval

* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs

* fix(benchmark): sub uri args; add report

* refactor: code format by ruff

* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf

* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search

* fix: adjust benchmark scripts

* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls

* refactor: new benchmark

* fix: step1 add resource by real code data

* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex

* optimize (benchmark): adjust keywords and ground truth for testing

* fix: truncate 64KB for content field

* optimize: effectiveness add resource plainly

* optimize: change param use of SearchByKeywords from "keywords" to "query"

* optimize(benchmark): refactor effectiveness scripts

* optimize: ensure raw data for content field

* optimize: fulltext analyzer's stop-words only use symbols

* fix: adapt to new ov cli for benchmark

* optimize: reuse file content to avoid re-read AGFS file

* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts

* optimize: benchmark client timeout

* update README

* fix: rm unused param

* fix: default values in docs

* optimize: increase truncate byte size to 1MB for content field for VikingDB

* fix(logger): harden queued stream logging (#2786)

* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock

When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.

During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.

Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.

Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.

Closes: #2752

* fix(logger): harden queued stream logging

---------

Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
2026-06-24 18:46:02 +08:00
Hao Zhe 324f96ebb6 fix(parse): normalize legacy text encodings (#2770)
* fix(parse): normalize text file encodings

* fix(parser): harden text encoding normalization

* fix(parse): normalize text encodings with charset-normalizer

* test(parse): use synthetic gb18030 fixture text

* fix(parse): respect detector rank for non-cjk text

* fix(parse): rescue short simplified chinese text

* fix(parse): preserve korean hanja text

* style(parse): format text encoding tests
2026-06-23 16:55:32 +08:00
AutoCoder ecced9930a feat(code-tools): add code navigation endpoints (#2671)
Add HTTP APIs for code outline, search, and expansion backed by the existing AST tooling.

Expose the same capabilities through the opencode plugin and cover the new routes, parser behavior, and plugin wiring with tests.
2026-06-17 16:12:19 +08:00
Qin Haojie 43a93d7ad9 feat(resource): 支持飞书用户 token 导入 (#2549)
* feat(resource): 支持飞书用户 token 导入

* feat(resource): 支持飞书用户 token watch

* fix(feishu): allow one-time user token imports

* test(feishu): trim redundant token coverage
2026-06-15 16:34:27 +08:00
8bd2b45d4c feat(parse): 目录/重入入库拆分图片并改写为 viking:// URI、Markdown 图片引用 viking 化 / split images and rewrite refs to viking:// URIs on directory & re-ingest (#2557)
[中文]
- 功能1:支持包含 PDF/DOC/Markdown 文件的目录入库后拆分图片并改写引用为 viking:// URI(此前 #2429 仅单文件入库生效,目录入库与重入场景下引用全部断链)
    - 图片相对引用此前只以 base_dir(md 自身目录)为根解析,../images/x.png 这类指向入库树内其他目录的引用解析失败 → DirectoryParser 把入库根 root_dir 作为 allowed_media_dirs 下传,_resolve_image_path 增加 markdown 语义(先相对 md 自身目录解析、落在任一允许根内即收)重新定位图片文件
    - 链接重写与图片入库的归属冲突:_rewrite_relative_links 此前跳过全部图片 embed,但图片入库只接管能解析且过校验的图片,越界/失效的两边都不管而断链 → _ingest_will_handle_image 探测分流(与入库同条件),会被接管的保持原样等 viking:// 改写、其余走相对路径深度调整
    - .image_mappings.json 边车在 _merge_temp/_recursive_move 合并 temp 树时被默认 ls(过滤隐藏文件)忽略导致下游无法消费 → 白名单模式(_MERGE_SIDECAR_ALLOWLIST)放行该边车拷贝,其余隐藏文件仍过滤
    - rewrite_image_uris 坐标系错位:此前只读资源根级 mapping 且按资源根路径查询,而 parser 是按各文档根写 mapping、key 相对文档根 → 改为探测每个 md 的祖先目录发现所有 mapping,以 mapping 所在目录为坐标系消费
    - 重入已存在资源(target_preexisting)走 SemanticProcessor 的 temp→target sync,该 sync 把可见文件 MOVE 进 target 并跳过隐藏文件,等改写时 temp 已无 md、边车搬不过去(rewrote 0 后被 delete_temp 销毁)→ 发现逻辑改由 target 树已 sync 的 md 驱动,取其祖先目录镜像回 temp 探测搬运边车
- 功能2:支持 Markdown 文件中图片相对引用路径改写成 viking uri 格式,支持两种格式 ![...](xxx) 和 <img src="xxx">,普通 markdown 链接语法(即使指向图片文件)未做改写
    - <img src="xxx"> 与 ![...] 全链路同等待遇(_ingest_local_images 采集、rewrite_image_uris 改写并保留 width 等其余属性、链接重写未接管时深度调整),正则经 image_rewrite.HTML_IMG_PATTERN 共享
- 其他(开发期间顺带):WordParser.parse 调用 parse_content 对齐 pdf.py 风格——删内层调用的 **kwargs 透传、显式转交 resource_name/source_name(原样 kwargs 值使命名逐字节不变)、allowed_media_dirs 简化为 [media_dir]
    - 顺手修复 #2429 引入的 test_word_parser_offloads_docx_conversion 回归:其 stub 未跟上 _convert_to_markdown 从 2 参改成 4 参

验收:用真实 ov add-resource 对隔离/真实服务端到端验证——目录入库 6/6 引用、docs/ 购买指南 3/3 个 <img> gif、情绪护肤 PDF 17/17、单文件 PDF 重入,全部改写为 viking:// URI、边车全部消费。tests/parse:380 通过,13 个既有失败(分支基线 cruft)保持不变;docx 资源命名由新增的 WordParser 转交测试锁定。

[EN]
- Feature 1: support image split + reference rewrite to viking:// URIs after directory ingest of trees containing PDF/DOC/Markdown files (#2429 only worked for single-file ingest; directory ingest and re-ingest left every reference broken)
    - Image relative refs were resolved against base_dir (the md's own dir) only, so ../images/x.png pointing elsewhere in the ingested tree failed to resolve → DirectoryParser passes the import root_dir down as allowed_media_dirs and _resolve_image_path gained markdown semantics (resolve relative to the md's own dir, accept if it stays inside ANY allowed root) to relocate the image
    - Ownership conflict between link rewriting and image ingestion: _rewrite_relative_links skipped ALL image embeds, but ingestion only takes images that resolve and pass validation, so out-of-scope/invalid ones were handled by neither and broke → _ingest_will_handle_image probes (same conditions as ingestion) and splits ownership: taken embeds stay untouched for the viking:// rewrite, the rest get relative-path depth adjustment
    - The .image_mappings.json sidecar was dropped by the default ls (hidden files filtered) during _merge_temp/_recursive_move, so downstream could not consume it → an allowlist (_MERGE_SIDECAR_ALLOWLIST) carries that sidecar while other hidden files stay filtered
    - rewrite_image_uris used the wrong coordinate system: it read only the resource-root mapping with root-relative keys, while the parser writes one mapping per document root with keys relative to it → discover every mapping by probing each md's ancestor directories and consume it in the coordinate system of the directory holding it
    - Re-ingesting an existing resource (target_preexisting) goes through SemanticProcessor's temp→target sync, which MOVES visible files into the target and skips hidden ones, so by rewrite time temp held no md and the sidecar couldn't be carried (rewrote 0, then delete_temp destroyed it) → discovery is now driven by the md files already synced into the target, mirroring their ancestor dirs back onto temp to probe and carry each sidecar
- Feature 2: rewrite Markdown image relative refs to viking uri, supporting both forms ![...](xxx) and <img src="xxx">; plain markdown link syntax (even when it targets an image file) is left unchanged
    - <img src="xxx"> gets the exact same pipeline treatment as ![...] (collected by _ingest_local_images, rewritten by rewrite_image_uris with width/other attributes preserved, depth-adjusted by link rewrite when not taken); pattern shared via image_rewrite.HTML_IMG_PATTERN
- Misc (incidental during development): align WordParser.parse's inner parse_content call with pdf.py — drop the **kwargs pass-through, forward resource_name/source_name explicitly (raw kwargs values so naming is byte-for-byte unchanged), and use [media_dir] as the only allowed_media_dirs
    - incidentally fixes the #2429-introduced test_word_parser_offloads_docx_conversion regression (its stub had not followed _convert_to_markdown's 2→4 arg change)

Verified end-to-end with real `ov add-resource` against isolated/real servers — directory ingest 6/6 references, docs/ purchase guide 3/3 <img> gifs, emotion-skincare PDF 17/17, single-file PDF re-ingest, all rewritten to viking:// URIs with sidecars consumed. tests/parse: 380 passed, 13 pre-existing failures (branch baseline cruft) unchanged; docx resource naming locked by a new WordParser forwarding test.

Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 15:42:18 +08:00
ac3c717b24 refactor(parse): split MarkdownParser.parse_content into parse + write phases / 重构 MarkdownParser.parse_content 为解析与写入两阶段 (#2529)
[EN]
1. Split parse_content() into two phases that can evolve independently:
   - _compute_layout(): parse only — turns the markdown into an ordered VikingFS
     write plan (list of _LayoutOp) holding raw section content, with zero side
     effects; the temp URI is passed in by the caller.
   - _apply_layout(): write only — replays the plan against VikingFS
     (_write_section, which does link rewrite, then _ingest_local_images).
   The relative-link probe now reuses the pure _compute_layout to learn a
   target's split layout instead of re-parsing through a fake filesystem, so
   _InMemoryFS is deleted entirely (and dead _try_add_to_pending /
   _flush_pending with it). _inmemory_split_files is renamed _probe_split_layout.

2. Link rewriting no longer handles image references: _rewrite_relative_links
   skips image embeds (![...]) and rewrites only document links. Image
   ingestion and image-path rewriting are owned entirely by _ingest_local_images.

[中文]
1. 将 parse_content() 拆成两个可独立演进的阶段:
   - _compute_layout():只解析——把 markdown 转成有序的 VikingFS 写入计划
     (_LayoutOp 列表,含原始 section 内容),零副作用;temp URI 由调用方传入。
   - _apply_layout():只写入——回放计划到 VikingFS(_write_section 内做链接重写,
     再 _ingest_local_images)。
   相对链接探测现在复用纯函数 _compute_layout 来获取目标的拆分布局,不再经假文件
   系统重新解析,故 _InMemoryFS 整类删除(连同死代码 _try_add_to_pending /
   _flush_pending)。_inmemory_split_files 改名为 _probe_split_layout。

2. 链接重写不再处理图片引用:_rewrite_relative_links 跳过图片 embed(![...]),
   只重写文档链接。图片入库与图片路径改写完全由 _ingest_local_images 负责。

Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 20:07:05 +08:00
Evo 5bc465a5f6 fix(parse/media): honor source_name/resource_name for media internal filenames (#2382) (#2492)
The image/audio/video parsers derived the resource's internal filename, viking
URI and title from file_path.name/file_path.stem -- i.e. the temp upload id
upload_<uuid> -- ignoring the caller-supplied resource_name/source_name, even
though media_processor already passes resource_name through and the markdown
parser honors it.

Factor the resolution into a shared resolve_media_names() helper used by all
three parsers: prefer resource_name (or source_name), stripping a trailing
suffix only when it is a known media extension (so "photo.png" doesn't double
its extension while "meeting.v1" is preserved); fall back to the temp file name
with byte-identical behavior to the previous inline logic. Unit tests cover the
helper across all cases (the three parsers share it) plus an image-parser
integration test. The extension still comes from the actual temp file.
2026-06-08 20:32:32 +08:00
bbed0c422a feat: save images from pdf and doc to vikingfs (#2429)
* feat: save images in Markdown to vikingfs

* feat: save images in doc to vikingfs

* fix: change the image file reading operation to asynchronous.

* feat: support generating L0 and L1 from images in Markdown and importing them into vector indexes

* fix: add updating image links when modifying existing files; restrict accessible paths for image loading; revise the unit tests raised for images

* fix: 限制保存的图片所在目录,避免将其他目录下的内容加入 vikingfs;对markdown代码中的图片链接保持不变

---------

Co-authored-by: zhanghaoyu.la <zhanghaoyu.la@bytedance.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-06-08 10:52:08 +08:00
960d04ed0b feat(parse): rewrite markdown relative links on directory ingest (#2452)
目录入库 markdown 时,MarkdownParser 会把每个 .md 按标题拆成目录结构,导致
[x](./other.md)、![](./img.png) 等相对链接失效——目标已不在原路径。本改动在写入
每个拆分 section 前重写相对链接,补偿入库引入的路径变换(源文件→目录、目标 .md→
目录或文件)。

When a markdown directory is ingested, MarkdownParser splits each .md into a
directory structure, breaking relative links like [x](./other.md) and
![](./img.png) — their targets no longer live at the original paths. This rewrites
relative links before writing each split section, compensating for the path
transformations ingest introduces.

核心设计 / Key properties:
- 磁盘坐标系 / Disk-coordinate: relpath 用原始磁盘路径计算,对 --to 免疫。
- 保守精确 / Conservative: 仅重写磁盘存在且落在 import_root 子树内的目标;外链 /
  页内锚点 / 绝对路径 / 缺失 / 越界一律原样保留。
- 对入库零假设 / Zero ingest assumptions: 目标 .md 的入库布局通过用同一个
  MarkdownParser 实跑解析到内存 FS 得到(_target_split_files/_inmemory_split_files);
  落点是文件还是目录(_doc_landing)、目录名、章节文件、章节内容全部来自 parser
  本身。无论 parse_content 今后怎么改,跑一遍 in-memory parse 即得最终结果,重写
  自动跟随、与真实入库一致,不复刻命名/拆分规则,也不假设 .md 一定目录化。

The target's ingest layout is obtained by running the SAME MarkdownParser into an
in-memory FS: whether the landing is a file or a directory, its name, the section
files and their content all come from the parser, so nothing about ingest is
reimplemented or assumed and the rewrite follows parse_content automatically.

带锚点的 .md 经 in-memory parse 精确定位到章节文件(GitHub-slug 匹配标题);单文件
文档(含未来小 .md 不拆目录)保留后缀指向该文件;图片 / 裸目录保持路径仅调相对深度。
重写链路为 async,且仅由 DirectoryParser 触发——单文件入库不重写。

新增 tests/parse/test_markdown_link_rewrite.py,22 passed。

Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 16:47:40 +08:00
Qin Haojie d606b59fb8 fix(parse): offload document conversions to thread (#2237)
Move synchronous document conversion work for Word, Excel, PowerPoint,
EPub, and legacy DOC parsers off the event loop, matching the PDF parser
behavior.
2026-05-26 12:02:04 +08:00
Rajvardhan PatilandRajvardhan Patil 1702815960 fix PDF XObject image extraction (#2199)
Co-authored-by: Rajvardhan Patil <243567420+RajvardhanPatil07@users.noreply.github.com>
2026-05-25 22:44:26 +08:00
Misaka 3d4bcc0a53 fix(parse): detect binary URL imports after GET (#2203) 2026-05-25 21:21:10 +08:00