Commit Graph
18 Commits
Author SHA1 Message Date
Eurakaxun 44c6df2622 perf: retrieval, import, LangChain, and session-context optimizations (#3569) 2026-07-30 10:10:11 +08:00
chenxiaobin-monkeyandchenxiaobin.monkey ff37e25cfd fix(parse): distinguish mpegts from TypeScript ts (#3574)
* fix(parse): distinguish mpegts from TypeScript ts

* fix(parse): tighten mpegts ts routing semantics

* fix(semantic): use file name for media summary type

---------

Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
2026-07-29 13:28:25 +08:00
Qin Haojie 2be4bb4879 refactor(parse): 收口资源解析路由 (#3295)
* refactor(parse): simplify resource ingestion routing

Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.

* fix(feishu): preserve sheet and bitable imports

Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.

* fix(parse): keep normalized Feishu content internal

Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.

* fix(feishu): parse bitable blocks embedded in sheets

Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.

* fix(feishu): download bitable attachment images

* refactor(parse): remove unused document converter

* refactor(parse): unify Understanding routing

* docs(parse): mark routing classification points

* docs(parse): complete wait routing flow

* fix(parse): preserve Feishu Base URL scope

* refactor(resource): separate ingestion submission from execution

* fix(resource): reject internal ingestion fields at public entry
2026-07-27 16:47:20 +08:00
MaojiaShengandqin-ctx 3c70f8d370 feat(parse): add large image processing for image parser (#3265)
* feat(parse): add large image processing for image parser

- Add large_image_processor.py: detect large images (>10MB or >4096px),
  create low-res previews, split into grid tiles, and generate grid
  overlay images with tile labels
- Refactor ImageParser.parse() to integrate large image processing pipeline
- Enable SVG-to-PNG conversion in utils.py (cairosvg/wand)
- Rename ImageConfig.max_dimension to preview_max_dimension and add new
  config fields: max_file_size_mb, max_tile_size_mb, max_tile_dimension_px,
  tile_overlap_px, large_image_threshold_dimension
- Update ov.conf.example with new image config options

* fix(parse): correct tile dimension comment from 1024px to 2048px

* fix(parse): fix tile label path in grid overlay to include tiles/ directory

* fix(parse): register missing image extensions for ImageParser

TIFF, ICO, DIB, ICNS, SGI, JP2 were not in IMAGE_EXTENSIONS, causing
them to fallback to TextParser. All are supported by PIL.

* fix(parse): preserve PNG format for tiles instead of always converting to JPEG

* fix(parse): address review feedback for large image processing

- Wire config.image to ImageParser in ParserRegistry (was missing)
- Remove unnecessary preview creation for small images (broke LA mode PNG)
- Enforce max_tile_size_mb on tiles with quality reduction and resize fallback
- Remove 64-tile hard cap that conflicted with max_tile_dimension_px
- Add comment explaining why original file is not saved for large images

* refactor(parse): remove max_tile_size_mb as it is a soft suggestion

max_tile_size_mb was a soft constraint that was not enforced
consistently. Remove it from config, constants, and all enforcement
logic. Tile dimension (max_tile_dimension_px) remains the sole constraint.

* fix(parse): use CJK-capable font for grid overlay labels

The old font loading only tried macOS-specific paths and fell back to
PIL's default bitmap font, which cannot render CJK characters in
filenames. Add a cross-platform CJK font lookup that covers Linux
(Noto/Droid/WQY/DejaVu), macOS (PingFang), and Windows (MSYH/SimSun).

* fix(parse): convert non-VLM-supported image formats to PNG on save

Image formats like TIFF, ICO, DIB, ICNS, SGI, JP2 are not recognized
by VLM backends (OpenAI/LiteLLM/VolcEngine only support PNG/JPEG/GIF/
WebP/BMP) or by embedding_utils for image vectorization. When a file
with one of these extensions is parsed, convert it to PNG and use a
.png extension so that downstream pipelines see consistent data.

SVG files (already PNG-converted via cairosvg) also get the .png
extension for the same reason.

* fix(parse): import io for SVG conversion

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-16 17:37:05 +08:00
Qin Haojie d47f2106ee refactor: remove unused and deprecated APIs (#3272)
Delete dead compatibility paths and test-only helpers so unsupported APIs do not remain as accidental contracts.
2026-07-16 10:49:56 +08:00
Jiahui Zhou 3bcefb298d fix: unify runtime loggers with openviking logger (#1981) 2026-05-12 11:26:30 +08:00
Jiahui Zhou 29ec5ca0c8 fix(parser): Fix parser config propagation for markdown splitting (#1480) 2026-04-15 22:50:25 +08:00
MaojiaSheng 95cc0f84d0 reorg: split parser layer to 2-layer: accessor and parser, so that we can reuse more code (#1428) 2026-04-15 10:33:12 +08:00
MaojiaShengandopenviking ce998873f9 lisence: change the main lisence to AGPL-3.0 (#1085)
* lisence: change the main lisence from Apache-2.0 to AGPL-v3

* lisence: change the main lisence from Apache-2.0 to AGPL-v3

* lisence: change the main lisence from Apache-2.0 to AGPL-v3

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-30 14:37:42 +08:00
ryzn 07cfc76cde feat(parse): add Feishu/Lark cloud document parser (#831)
Add FeishuParser that supports importing Feishu cloud documents
(docx, wiki, sheets, bitable) into the knowledge base via URL.

Supports:
- Docx documents via Blocks API with attribute-driven block detection
- Wiki pages (auto-resolve to underlying document type)
- Spreadsheets via Sheets API
- Bitable (multi-dimensional tables) via Bitable API
- Embedded sheet views inside docx documents
- Generic text extraction fallback for unknown block types

Design:
- Uses lark-oapi SDK for all API calls (auth, pagination, etc.)
- Attribute-driven block detection: inspects which SDK attribute is
  populated rather than hard-coding 50+ block type integer constants
- Follows convert-then-parse pattern: Feishu -> Markdown -> MarkdownParser
- Lazy imports to avoid breaking when lark-oapi is not installed
- FeishuConfig for credentials (env vars or ov.conf)

Integration:
- URL routing in media_processor.py for feishu.cn/larksuite.com
- Parser registration with ImportError guard
- parse-feishu optional dependency group in pyproject.toml

Tested end-to-end: Feishu URL -> parse -> VikingFS -> L0/L1 generation
-> vectorization -> semantic search, all working.
2026-03-22 11:56:53 +08:00
Lam Ngoc Nguyen 559eef38c9 feat(parse): add support for legacy .doc and .xls file formats (#652)
* feat(parse): add support for legacy .doc and .xls file formats

Add LegacyDocParser using olefile to extract text from Word 97-2003
binary .doc files via OLE2 stream parsing with piece table support
and multi-level fallbacks.

Extend ExcelParser to handle .xls files using xlrd, branching the
parse logic based on file extension while keeping openpyxl for
.xlsx/.xlsm.

New dependencies: olefile>=0.47, xlrd>=2.0.1

* fix(parse): handle date and boolean cell types in xlrd .xls parsing

Check cell.ctype for XL_CELL_DATE and XL_CELL_BOOLEAN to avoid
outputting raw float serial numbers for dates and numeric 0/1 for
booleans.

* fix(parse): harden legacy .doc/.xls parsers per code review

excel.py:
- Enable formatting_info=True so xlrd detects date cells properly
- Add on_demand=True and release_resources() for memory efficiency
- Handle all xlrd cell types: DATE (with time), BOOLEAN, ERROR, BLANK, EMPTY
- Display integers without trailing .0
- Extract cell formatting to _format_xls_cell static method

legacy_doc.py:
- Add 50MB stream size cap to prevent DoS from crafted files
- Cap ccpText at 10M chars to prevent memory exhaustion
- Add FIB version check (require Word 97+ / nFib >= 0x00C1)
- Add minimum buffer length check before struct.unpack_from
- Fix Grpprl skip loop to prevent spin on zero-length entries
- Add _clean_word_text for \x0B (soft break) and \x0C (section break)
- Log warnings for pieces extending beyond stream bounds
- Cap fallback extract to max stream size
2026-03-16 16:29:51 +08:00
Eric Shaw 6a9cd20a67 tests(parsers): add unit tests for office extensions within add_resource directory (#273) 2026-02-25 13:00:55 +08:00
MaojiaShengandopenviking dc7bc956f3 feat: update media parsers (#196)
* fix: make rust CLI (ov) commands match python CLI (openviking) exactly - add top-level wait/status/health commands

* fix: ov cli plays same as py cli (ls, tree)

* Refactor media parsers to subdirectory structure with validation

* Enhance CLI robustness: validate add-resource path exists and detect unquoted spaces

* Fix unescaped spaces in paths by replacing \  with space

* Sanitize URI components to replace spaces and special chars with underscores

* feat: auto organize audio and image and video files

* Update media parsers to use original filenames and folder names with extensions

* Optimize MediaParser section for readability

* feat: vlm optimization for image

* feat: vlm optimization for image

* feat: vlm optimization for image

* refactor: move media content understanding to SemanticProcessor

- Add parse/parsers/media/utils.py with media helpers
- Refactor ImageParser.parse(), AudioParser.parse(), VideoParser.parse() to remove content understanding, keep only metadata extraction
- Update SemanticProcessor._generate_single_file_summary() to handle media types and call media utils for summary generation
- Update TreeBuilder._get_base_uri() to use media utils
- Update ResourceNode.get_abstract() and get_overview() to check meta for abstract/overview
- Add debug logs and error handling

* refactor: split _generate_single_file_summary to add _generate_text_summary

- Add _generate_text_summary function for text file processing
- Update media utils functions to accept llm_sem and use it to limit concurrent calls
- Update _generate_single_file_summary to call _generate_text_summary and media utils functions
- Fix import ordering
- Fix issue where _generate_file_summaries was creating a new semaphore, now each _generate_single_file_summary handles its own

* feat: vlm optimization for image

* feat: vlm optimization for image

---------

Co-authored-by: openviking <openviking@example.com>
2026-02-20 15:42:19 +08:00
Eric Shaw 780e36a39c feat: add directory parsing support to OpenViking (#194)
* feat: add directory parsing support to OpenViking

- Implemented DirectoryParser to handle local directories with mixed document types.
- Enhanced add_resource function to support directory imports with options for including, excluding, and ignoring specific directories.
- Updated client and service layers to forward additional parsing options.
- Added unit tests for DirectoryParser to ensure correct functionality and error handling.
- Improved user feedback with rich table summaries for processed, failed, unsupported, and skipped files during directory imports.

* docs: update README.md to include directory import instructions for add.py

* style: reformat files to pass CI code formatting

* style: reformat files to pass CI code formatting
2026-02-16 16:21:08 +08:00
Zayn Jarvis c3042fec11 feat: add markitdown-inspired file parsers (Word, PowerPoint, Excel, EPub, ZIP) (#128)
* feat: add markitdown-inspired file parsers

Add support for parsing additional file formats inspired by microsoft/markitdown:

- Word (.docx) - using python-docx
- PowerPoint (.pptx) - using python-pptx
- Excel (.xlsx) - using openpyxl
- Audio (.mp3, .wav, .m4a, etc.) - metadata extraction using mutagen
- EPub (.epub) - using ebooklib
- ZIP (.zip) - iterate contents

All parsers convert content to markdown and delegate to MarkdownParser
for tree structure creation, following OpenViking's parser pattern.

Dependencies added to pyproject.toml:
- python-docx, python-pptx, openpyxl
- ebooklib, beautifulsoup4
- mutagen

Includes comprehensive tests for all new parsers.

Refs: markitdown-parsers

* feat: make markitdown parsers built-in capabilities

Move parser dependencies from optional to main dependencies.
Register parsers directly without graceful fallback.
Remove optional registration infrastructure.

Parsers now built-in:
- Word (.docx) via python-docx
- PowerPoint (.pptx) via python-pptx
- Excel (.xlsx) via openpyxl
- EPub (.epub) via ebooklib
- ZIP (.zip) via built-in zipfile
- Audio (.mp3, .wav, etc.) via mutagen

* refactor: align parsers with existing ecosystem

Set source_format on ParseResult like TextParser does:
- word: source_format = 'docx'
- powerpoint: source_format = 'pptx'
- excel: source_format = 'xlsx'
- epub: source_format = 'epub'
- zip: source_format = 'zip'
- audio: source_format = 'audio'

All parsers now follow the same pattern as existing TextParser
and PDFParser for consistency.

* fix: align markitdown parsers with ecosystem patterns

Critical fixes:
- Remove duplicate zip_archive.py (conflicting ZipParser class name)
- Use zip_parser.py as canonical ZIP parser (follows TextParser pattern)
- Fix parse_content() to delegate to MarkdownParser instead of raising
  ValueError (all parse_content tests were broken)
- Set parser_name on all ParseResult outputs (was missing)
- Set source_format AFTER MarkdownParser call (was being overwritten)
- Accept ParserConfig in all parser __init__ (ecosystem consistency)
- Add .xlsm to ExcelParser supported_extensions
- Fix AudioParser._format_size to match ZipParser format (500.0 B)
- Fix pyproject.toml urllib3 indentation corruption
- Add tests/parse/conftest.py with VikingFS test fixture
- Rewrite tests to actually pass and cover registry integration

All 23 tests passing.

* fix: improve markitdown parser consistency and add real-file tests

Critical fixes:
- WordParser: preserve table position in document order (was appending
  all tables at end, losing context). Walk document body XML in order
  instead of iterating paragraphs then tables separately.
- PowerPointParser: replace magic number (type == 1) with proper
  PP_PLACEHOLDER enum constants, also handle CENTER_TITLE.
- AudioParser: add Vorbis/FLAC/OGG tag extraction (previously only
  handled ID3 and MP4 formats). Tries all format mappings with dedup.
- ZipParser: replace emoji in tree view with plain text markers
  for robustness in text processing pipelines.
- TextParser: set parser_name='TextParser' on parse_content results
  for consistency with all other parsers.
- __init__.py: export all new parser classes for public API.

Tests (16 new, 39 total):
- Real .docx/.xlsx/.pptx file creation and parsing
- EPub HTML-to-markdown conversion edge cases
- ZIP bad-file error handling and no-emoji tree view
- AudioParser Vorbis tag extraction and edge cases
- WordParser can_parse() extension matching

* feat: update for rebase and remove audio redundant

* feat: rollback unexpected change

* chore: remove redundant mutagen for audio file
2026-02-14 16:55:26 +08:00
MaojiaSheng a9b59cce87 support small github code repos (#70)
* feat: add code parser, not tested

* feat: rename CodeParser to CodeRepositoryParser

* fix: code parser data uploading

* fix: code parser data uploading

* fix: code parsing

* feat: protect binary files to summary and vectorize

* Update viking_fs.py
2026-02-05 19:18:10 +08:00
kkkwjx a30cb2e7b9 Lint code (#25)
* feat: pre-commit-config

* style: ruff code
2026-02-02 18:56:15 +08:00
qin-ctx f98dc0ed1c first commit 2026-01-29 20:29:19 +08:00