Files
OpenViking/openviking/parse
yangxinxin-7andClaude Opus 5 8b4deaab99 fix(parser): drop duplicate images when a PDF stacks XObjects on one spot (#3662)
Print-to-PDF producers routinely emit several image XObjects drawn at the
exact same position on a page (a background layer plus a content layer).
Because `_extract_image_from_page` rasterises the page *region* rather than
decoding the XObject itself, every one of them renders to identical bytes —
so a document with two stacked full-page layers wrote two byte-identical
PNGs per page and referenced both from the generated markdown.

Dedup within each page, in two steps:

- bbox first, so a repeat is skipped before paying for the render;
- a content hash as a backstop, for bboxes that differ slightly but still
  rasterise to the same bytes.

Both sets are per-page, so a header logo repeated across pages is still
kept once on every page. `meta["images_deduplicated"]` reports how many
were skipped.

Measured on an 8-page article exported from a web page: 16 saved PNGs -> 8,
16 markdown image references -> 8, local conversion 3.8s -> 2.4s.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:03:21 +08:00
..