* fix(embedding): downsample oversized image inputs
Keep imported image resources unchanged while avoiding provider-side multimodal embedding failures for oversized images. The embedding path now builds a temporary downsampled image data URI when image bytes exceed the shared large-image limits.
Move reusable image size thresholds into media_limits so both parser-side large image handling and embedding-side input preparation depend on a common utility instead of embedding_utils importing parser internals.
Add vectorize_file coverage confirming large image embedding inputs are resized and the stored resource bytes are preserved.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(media): downsample oversized image model inputs
Keep imported image resources unchanged while avoiding provider-side multimodal failures for oversized images. Shared image input preparation now builds temporary downsampled bytes for model requests when image bytes exceed the configured large-image limits.
Apply the model-input downsampling to both semantic image summary generation and embedding image data URI construction, so directory and code repository imports can preserve original images while sending provider-compatible inputs.
Move reusable image size thresholds into media_limits so parser-side large image handling, VLM summary generation, and embedding preparation share common limits without embedding_utils importing parser internals.
Always convert downsampled model images to RGB before JPEG encoding so Pillow-openable modes such as LA or I;16 do not fall back to the original oversized bytes.
Add coverage confirming VLM image summaries, vectorize_file embedding inputs, and JPEG-incompatible image modes are resized while stored resource bytes are preserved.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Frontmatter was parsed into ParseResult.meta and removed from the body, but
that metadata is never persisted, so every ingested markdown file lost its
frontmatter with no way to read the fields back.
Parse frontmatter into meta unconditionally (it still drives doc_title) and
only remove it from the stored body when explicitly configured; that removal
is now off by default.
_smart_split_content documents that it enforces both a token limit
(max_size) and a hard character limit. But when a single paragraph was
oversized by tokens yet under the character limit, the force-split loop
stepped through it by max_chars only, so it emitted a chunk that still
exceeded max_size tokens.
This is reachable with ordinary long-form CJK text: _estimate_token_count
weights CJK at ~0.7 token/char, so a ~5000-char Chinese paragraph is
~3500 tokens (over the 2048 default) while staying under the char limit,
and was returned as a single over-budget chunk.
Bound the force-split step by min(max_chars, max_size / MAX_TOKENS_PER_CHAR),
where MAX_TOKENS_PER_CHAR is the worst-case (CJK) density already used by
_estimate_token_count, now extracted into a shared constant so the two stay
in sync.
Add regression tests for the token budget and content preservation.
_gh_slug claims to produce "GitHub-style" heading slugs but collapsed runs
of whitespace (`re.sub(r"\s+", "-", s)`). GitHub's reference slugger
(github-slugger) maps each space to its own hyphen (`.replace(/ /g, '-')`)
and does not collapse.
Because punctuation is stripped before spaces are converted, a heading like
"Foo & Bar" leaves two adjacent spaces where "&" was. GitHub renders this as
"foo--bar", but _gh_slug produced "foo-bar". The intra-document link rewriter
(_rewrite_link) compares _gh_slug(heading) against the link fragment, so an
author-written link such as `guide.md#foo--bar` failed to match its heading
and was left unrewritten after the target doc was split into sections.
Replace `\s+` with `\s` so each whitespace character maps to one hyphen,
matching GitHub. Simple single-space headings are unaffected.
Add regression tests covering punctuation headings and ordinary headings.
* feat: add audio and video understanding via VLM
* docs: design media resource guards
* fix: bound media staging concurrency
* fix: cap unknown-size media staging
* test: stage media in routing fake
* test: exercise media staging callbacks
* test: trim media understanding coverage
* chore: 清理实现计划文档
* fix: 修复多凭证切换问题
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Print-to-PDF producers routinely emit several image XObjects drawn at the
exact same position on a page (a background layer plus a content layer).
Because `_extract_image_from_page` rasterises the page *region* rather than
decoding the XObject itself, every one of them renders to identical bytes —
so a document with two stacked full-page layers wrote two byte-identical
PNGs per page and referenced both from the generated markdown.
Dedup within each page, in two steps:
- bbox first, so a repeat is skipped before paying for the render;
- a content hash as a backstop, for bboxes that differ slightly but still
rasterise to the same bytes.
Both sets are per-page, so a header logo repeated across pages is still
kept once on every page. `meta["images_deduplicated"]` reports how many
were skipped.
Measured on an 8-page article exported from a web page: 16 saved PNGs -> 8,
16 markdown image references -> 8, local conversion 3.8s -> 2.4s.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(parse): distinguish mpegts from TypeScript ts
* fix(parse): tighten mpegts ts routing semantics
* fix(semantic): use file name for media summary type
---------
Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
* refactor(parse): simplify resource ingestion routing
Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.
* fix(feishu): preserve sheet and bitable imports
Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.
* fix(parse): keep normalized Feishu content internal
Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.
* fix(feishu): parse bitable blocks embedded in sheets
Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.
* fix(feishu): download bitable attachment images
* refactor(parse): remove unused document converter
* refactor(parse): unify Understanding routing
* docs(parse): mark routing classification points
* docs(parse): complete wait routing flow
* fix(parse): preserve Feishu Base URL scope
* refactor(resource): separate ingestion submission from execution
* fix(resource): reject internal ingestion fields at public entry
中文:对 HTTP/HTTPS URL 使用 urlparse(url).path 提取扩展名,确保 video.mp4?signature=... 命中 UnderstandingAPI fast path。新增带签名视频 URL 的 ParserRouter 回归测试。
English: Parse HTTP/HTTPS URL paths before checking extensions so signed video URLs hit the UnderstandingAPI fast path. Add a ParserRouter regression test for signed video URLs.
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
* fix(feishu): surface permission errors clearly and keep users on page
Map Feishu/Lark API failures to typed OpenViking errors with actionable hints, and keep Web Studio from treating HTTP 403 permission denials as session logout.
* fix(feishu): simplify API error mapping
* refactor(feishu): inline API error mapping
---------
Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(feishu): persist imported document images
Download Feishu image tokens into local import temp trees so Markdown image references can be ingested alongside the document.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(feishu): harden inline-image download (async, extension, user token)
Address review feedback on the inline-image import path:
- Run the synchronous lark-oapi media download via asyncio.to_thread so a
slow Feishu request no longer blocks unrelated async work on the event loop,
matching the existing _fetch_document() pattern.
- Infer the image file extension from the downloaded bytes (byte-magic
sniffing) and fall back to the response Content-Type via the existing
mime_types.get_preferred_extension helper, instead of hardcoding .png. This
stops JPEG/WebP/GIF bytes from being mislabeled as PNG to downstream
consumers (e.g. the data:image/... URI built during multimodal vectorization).
- Advertise AccessTokenType.USER on the media download request when a user
access token is supplied, so lark-oapi actually injects it. Previously the
request only allowed TENANT, so user-token imports read the document body
but silently dropped every image.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(storage): optimize glob func
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* fix(localfs): offload blocking fs operations to spawn_blocking
* feat(glob): cap glob api default node_limit at 256
* feat(sdk): add node_limit options for glob in python and go SDKs
* refactor(web-crawler): remove Playwright rendering, keep static SPA shells
Drop the Playwright-based fallback rendering path so SPA pages are stored
as their static shell HTML instead of being rendered headlessly. SPA
shells now surface the <noscript> notice (e.g. "You need to enable
JavaScript to run this app.") rather than producing an empty document,
and robots.txt-blocked entry pages get a human-readable error message.
- Delete playwright_renderer.py, render_heuristics.py and their tests
- Strip fallback_playwright/playwright_timeout config, fallback_rendered
counter, and render-hint plumbing from the crawler and web importer
- HTMLParser: drop the SPA-empty-pattern stripping and fall back to
<noscript> text when trafilatura extracts nothing
- Humanize the robots.txt entry-failure message
- Migrate the orphaned _convert_to_raw_url tests to HTTPAccessor (the
method moved there in an earlier reorg) and remove the stale file
* refactor(web-importer): simplify robots.txt-blocked import message
Replace the verbose robots.txt explanation with a short compliance hint
that omits the URL and points users to local-file import instead.
* Refactor recursive web import into HTTP accessor
Move ordinary web page import routing into HTTPAccessor and materialize crawled pages as a temporary directory via WebImporter.
Relocate Scrapy/Playwright crawling under parse.accessors.web_crawler, keep trafilatura extraction inside HTMLParser, and avoid repeated ResourceService.add_resource calls.
Add recursive crawl controls, safe request validation, page/download classification, and focused unit coverage.
* Document recursive web crawler options
* fix(web-crawler): stop SSRF sub-resource block from failing whole render
The playwright fallback validated every sub-resource request against the
SSRF guard and raised on the first disallowed host, failing the entire
page render. volcengine docs load a probe resource on an internal host,
so rendering always failed and the crawler stored the static anti-bot
"Please wait..." challenge page as content.
Now a blocked sub-resource is only aborted; the main document and final
URL still gate the result. Also wait past JS interstitials, retry reads
through in-flight navigation, and reject shell/challenge pages instead of
storing them.
* fix(web-crawler): surface renderer error hint on entry-page failure
When Playwright is unavailable, the renderer returns an actionable install
hint via RenderResult.error, but the spider silently kept the static shell
and WebImporter raised only the generic "Failed to fetch entry page". The
hint never reached the user.
Now the spider records rendered.error on the failed page, and WebImporter
appends the entry page's failure reason to the raised message so the CLI
shows the Playwright install instructions.
* fix(web-crawler): surface render hints and enforce crawl limits
* fix(web-crawler): avoid rendering SSR app pages
* perf(web-crawler): bound render concurrency and cap networkidle wait
Playwright renders were dispatched from parse callbacks without any
concurrency limit, so a page with many child links could spawn dozens of
Chromium pages at once (observed peak 28 for a 20-page crawl), risking OOM
on large sites and starting ~2.3x more renders than needed before
max_pages stopped the crawl. Gate renders with a semaphore sized to
config.concurrency and re-check the success limit after acquiring a slot
so queued callbacks skip rendering once the crawl is already done.
Also cap the networkidle wait at 8s: pages with continuous background
activity (e.g. GraphiQL) never go idle and previously blocked until the
full render timeout, turning a ~3s page into ~38s. Content is ready after
domcontentloaded and _wait_past_challenge covers late-arriving text.
Bump default concurrency 5 -> 10.
* fix(web-crawler): route .html/.htm URLs through recursive WebImporter
An explicit .html/.htm URL is detected as DOWNLOAD_HTML via the extension
map, so access() previously only routed URLType.WEBPAGE to WebImporter and
these URLs fell through to single-file download, silently ignoring
depth/max_pages. Route DOWNLOAD_HTML through WebImporter too, treating a
single-page import as the depth=0 case.
* fix(web-crawler): improve HTML extraction and rendering heuristics
- Drop trafilatura favor_precision=True: it stripped the full body of
link-dense pages, keeping only headers.
- Only render __NEXT_DATA__ pages with Playwright when their static body
is too thin; SSR/SSG Next.js pages already ship full text.
- Disable Scrapy telnet console to avoid opening port 6023.
* fix(web-crawler): keep code-hosting single-file URLs off recursive crawler
GitHub/GitLab blob and GitHub raw URLs resolve to a single file, not a
site. Route them through the single-file download path instead of the
recursive WebImporter, which otherwise crawls the hosting UI shell.
* docs(resources): add recursive web crawler usage examples
Add depth/max_pages crawl examples to the HTTP, Python SDK, and CLI
blocks in both the zh and en resource API docs, plus path-prefix
filtering and skip_download_links variants.
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* fix(parse): normalize text file encodings
* fix(parser): harden text encoding normalization
* fix(parse): normalize text encodings with charset-normalizer
* test(parse): use synthetic gb18030 fixture text
* fix(parse): respect detector rank for non-cjk text
* fix(parse): rescue short simplified chinese text
* fix(parse): preserve korean hanja text
* style(parse): format text encoding tests
Add HTTP APIs for code outline, search, and expansion backed by the existing AST tooling.
Expose the same capabilities through the opencode plugin and cover the new routes, parser behavior, and plugin wiring with tests.
[中文]
- 功能1:支持包含 PDF/DOC/Markdown 文件的目录入库后拆分图片并改写引用为 viking:// URI(此前 #2429 仅单文件入库生效,目录入库与重入场景下引用全部断链)
- 图片相对引用此前只以 base_dir(md 自身目录)为根解析,../images/x.png 这类指向入库树内其他目录的引用解析失败 → DirectoryParser 把入库根 root_dir 作为 allowed_media_dirs 下传,_resolve_image_path 增加 markdown 语义(先相对 md 自身目录解析、落在任一允许根内即收)重新定位图片文件
- 链接重写与图片入库的归属冲突:_rewrite_relative_links 此前跳过全部图片 embed,但图片入库只接管能解析且过校验的图片,越界/失效的两边都不管而断链 → _ingest_will_handle_image 探测分流(与入库同条件),会被接管的保持原样等 viking:// 改写、其余走相对路径深度调整
- .image_mappings.json 边车在 _merge_temp/_recursive_move 合并 temp 树时被默认 ls(过滤隐藏文件)忽略导致下游无法消费 → 白名单模式(_MERGE_SIDECAR_ALLOWLIST)放行该边车拷贝,其余隐藏文件仍过滤
- rewrite_image_uris 坐标系错位:此前只读资源根级 mapping 且按资源根路径查询,而 parser 是按各文档根写 mapping、key 相对文档根 → 改为探测每个 md 的祖先目录发现所有 mapping,以 mapping 所在目录为坐标系消费
- 重入已存在资源(target_preexisting)走 SemanticProcessor 的 temp→target sync,该 sync 把可见文件 MOVE 进 target 并跳过隐藏文件,等改写时 temp 已无 md、边车搬不过去(rewrote 0 后被 delete_temp 销毁)→ 发现逻辑改由 target 树已 sync 的 md 驱动,取其祖先目录镜像回 temp 探测搬运边车
- 功能2:支持 Markdown 文件中图片相对引用路径改写成 viking uri 格式,支持两种格式  和 <img src="xxx">,普通 markdown 链接语法(即使指向图片文件)未做改写
- <img src="xxx"> 与 ![...] 全链路同等待遇(_ingest_local_images 采集、rewrite_image_uris 改写并保留 width 等其余属性、链接重写未接管时深度调整),正则经 image_rewrite.HTML_IMG_PATTERN 共享
- 其他(开发期间顺带):WordParser.parse 调用 parse_content 对齐 pdf.py 风格——删内层调用的 **kwargs 透传、显式转交 resource_name/source_name(原样 kwargs 值使命名逐字节不变)、allowed_media_dirs 简化为 [media_dir]
- 顺手修复 #2429 引入的 test_word_parser_offloads_docx_conversion 回归:其 stub 未跟上 _convert_to_markdown 从 2 参改成 4 参
验收:用真实 ov add-resource 对隔离/真实服务端到端验证——目录入库 6/6 引用、docs/ 购买指南 3/3 个 <img> gif、情绪护肤 PDF 17/17、单文件 PDF 重入,全部改写为 viking:// URI、边车全部消费。tests/parse:380 通过,13 个既有失败(分支基线 cruft)保持不变;docx 资源命名由新增的 WordParser 转交测试锁定。
[EN]
- Feature 1: support image split + reference rewrite to viking:// URIs after directory ingest of trees containing PDF/DOC/Markdown files (#2429 only worked for single-file ingest; directory ingest and re-ingest left every reference broken)
- Image relative refs were resolved against base_dir (the md's own dir) only, so ../images/x.png pointing elsewhere in the ingested tree failed to resolve → DirectoryParser passes the import root_dir down as allowed_media_dirs and _resolve_image_path gained markdown semantics (resolve relative to the md's own dir, accept if it stays inside ANY allowed root) to relocate the image
- Ownership conflict between link rewriting and image ingestion: _rewrite_relative_links skipped ALL image embeds, but ingestion only takes images that resolve and pass validation, so out-of-scope/invalid ones were handled by neither and broke → _ingest_will_handle_image probes (same conditions as ingestion) and splits ownership: taken embeds stay untouched for the viking:// rewrite, the rest get relative-path depth adjustment
- The .image_mappings.json sidecar was dropped by the default ls (hidden files filtered) during _merge_temp/_recursive_move, so downstream could not consume it → an allowlist (_MERGE_SIDECAR_ALLOWLIST) carries that sidecar while other hidden files stay filtered
- rewrite_image_uris used the wrong coordinate system: it read only the resource-root mapping with root-relative keys, while the parser writes one mapping per document root with keys relative to it → discover every mapping by probing each md's ancestor directories and consume it in the coordinate system of the directory holding it
- Re-ingesting an existing resource (target_preexisting) goes through SemanticProcessor's temp→target sync, which MOVES visible files into the target and skips hidden ones, so by rewrite time temp held no md and the sidecar couldn't be carried (rewrote 0, then delete_temp destroyed it) → discovery is now driven by the md files already synced into the target, mirroring their ancestor dirs back onto temp to probe and carry each sidecar
- Feature 2: rewrite Markdown image relative refs to viking uri, supporting both forms  and <img src="xxx">; plain markdown link syntax (even when it targets an image file) is left unchanged
- <img src="xxx"> gets the exact same pipeline treatment as ![...] (collected by _ingest_local_images, rewritten by rewrite_image_uris with width/other attributes preserved, depth-adjusted by link rewrite when not taken); pattern shared via image_rewrite.HTML_IMG_PATTERN
- Misc (incidental during development): align WordParser.parse's inner parse_content call with pdf.py — drop the **kwargs pass-through, forward resource_name/source_name explicitly (raw kwargs values so naming is byte-for-byte unchanged), and use [media_dir] as the only allowed_media_dirs
- incidentally fixes the #2429-introduced test_word_parser_offloads_docx_conversion regression (its stub had not followed _convert_to_markdown's 2→4 arg change)
Verified end-to-end with real `ov add-resource` against isolated/real servers — directory ingest 6/6 references, docs/ purchase guide 3/3 <img> gifs, emotion-skincare PDF 17/17, single-file PDF re-ingest, all rewritten to viking:// URIs with sidecars consumed. tests/parse: 380 passed, 13 pre-existing failures (branch baseline cruft) unchanged; docx resource naming locked by a new WordParser forwarding test.
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
[EN]
1. Split parse_content() into two phases that can evolve independently:
- _compute_layout(): parse only — turns the markdown into an ordered VikingFS
write plan (list of _LayoutOp) holding raw section content, with zero side
effects; the temp URI is passed in by the caller.
- _apply_layout(): write only — replays the plan against VikingFS
(_write_section, which does link rewrite, then _ingest_local_images).
The relative-link probe now reuses the pure _compute_layout to learn a
target's split layout instead of re-parsing through a fake filesystem, so
_InMemoryFS is deleted entirely (and dead _try_add_to_pending /
_flush_pending with it). _inmemory_split_files is renamed _probe_split_layout.
2. Link rewriting no longer handles image references: _rewrite_relative_links
skips image embeds (![...]) and rewrites only document links. Image
ingestion and image-path rewriting are owned entirely by _ingest_local_images.
[中文]
1. 将 parse_content() 拆成两个可独立演进的阶段:
- _compute_layout():只解析——把 markdown 转成有序的 VikingFS 写入计划
(_LayoutOp 列表,含原始 section 内容),零副作用;temp URI 由调用方传入。
- _apply_layout():只写入——回放计划到 VikingFS(_write_section 内做链接重写,
再 _ingest_local_images)。
相对链接探测现在复用纯函数 _compute_layout 来获取目标的拆分布局,不再经假文件
系统重新解析,故 _InMemoryFS 整类删除(连同死代码 _try_add_to_pending /
_flush_pending)。_inmemory_split_files 改名为 _probe_split_layout。
2. 链接重写不再处理图片引用:_rewrite_relative_links 跳过图片 embed(![...]),
只重写文档链接。图片入库与图片路径改写完全由 _ingest_local_images 负责。
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The image/audio/video parsers derived the resource's internal filename, viking
URI and title from file_path.name/file_path.stem -- i.e. the temp upload id
upload_<uuid> -- ignoring the caller-supplied resource_name/source_name, even
though media_processor already passes resource_name through and the markdown
parser honors it.
Factor the resolution into a shared resolve_media_names() helper used by all
three parsers: prefer resource_name (or source_name), stripping a trailing
suffix only when it is a known media extension (so "photo.png" doesn't double
its extension while "meeting.v1" is preserved); fall back to the temp file name
with byte-identical behavior to the previous inline logic. Unit tests cover the
helper across all cases (the three parsers share it) plus an image-parser
integration test. The extension still comes from the actual temp file.
* feat: save images in Markdown to vikingfs
* feat: save images in doc to vikingfs
* fix: change the image file reading operation to asynchronous.
* feat: support generating L0 and L1 from images in Markdown and importing them into vector indexes
* fix: add updating image links when modifying existing files; restrict accessible paths for image loading; revise the unit tests raised for images
* fix: 限制保存的图片所在目录,避免将其他目录下的内容加入 vikingfs;对markdown代码中的图片链接保持不变
---------
Co-authored-by: zhanghaoyu.la <zhanghaoyu.la@bytedance.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
目录入库 markdown 时,MarkdownParser 会把每个 .md 按标题拆成目录结构,导致
[x](./other.md)、 等相对链接失效——目标已不在原路径。本改动在写入
每个拆分 section 前重写相对链接,补偿入库引入的路径变换(源文件→目录、目标 .md→
目录或文件)。
When a markdown directory is ingested, MarkdownParser splits each .md into a
directory structure, breaking relative links like [x](./other.md) and
 — their targets no longer live at the original paths. This rewrites
relative links before writing each split section, compensating for the path
transformations ingest introduces.
核心设计 / Key properties:
- 磁盘坐标系 / Disk-coordinate: relpath 用原始磁盘路径计算,对 --to 免疫。
- 保守精确 / Conservative: 仅重写磁盘存在且落在 import_root 子树内的目标;外链 /
页内锚点 / 绝对路径 / 缺失 / 越界一律原样保留。
- 对入库零假设 / Zero ingest assumptions: 目标 .md 的入库布局通过用同一个
MarkdownParser 实跑解析到内存 FS 得到(_target_split_files/_inmemory_split_files);
落点是文件还是目录(_doc_landing)、目录名、章节文件、章节内容全部来自 parser
本身。无论 parse_content 今后怎么改,跑一遍 in-memory parse 即得最终结果,重写
自动跟随、与真实入库一致,不复刻命名/拆分规则,也不假设 .md 一定目录化。
The target's ingest layout is obtained by running the SAME MarkdownParser into an
in-memory FS: whether the landing is a file or a directory, its name, the section
files and their content all come from the parser, so nothing about ingest is
reimplemented or assumed and the rewrite follows parse_content automatically.
带锚点的 .md 经 in-memory parse 精确定位到章节文件(GitHub-slug 匹配标题);单文件
文档(含未来小 .md 不拆目录)保留后缀指向该文件;图片 / 裸目录保持路径仅调相对深度。
重写链路为 async,且仅由 DirectoryParser 触发——单文件入库不重写。
新增 tests/parse/test_markdown_link_rewrite.py,22 passed。
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Move synchronous document conversion work for Word, Excel, PowerPoint,
EPub, and legacy DOC parsers off the event loop, matching the PDF parser
behavior.