* fix(parse): distinguish mpegts from TypeScript ts
* fix(parse): tighten mpegts ts routing semantics
* fix(semantic): use file name for media summary type
---------
Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
* refactor(parse): simplify resource ingestion routing
Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.
* fix(feishu): preserve sheet and bitable imports
Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.
* fix(parse): keep normalized Feishu content internal
Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.
* fix(feishu): parse bitable blocks embedded in sheets
Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.
* fix(feishu): download bitable attachment images
* refactor(parse): remove unused document converter
* refactor(parse): unify Understanding routing
* docs(parse): mark routing classification points
* docs(parse): complete wait routing flow
* fix(parse): preserve Feishu Base URL scope
* refactor(resource): separate ingestion submission from execution
* fix(resource): reject internal ingestion fields at public entry
* feat(parse): add large image processing for image parser
- Add large_image_processor.py: detect large images (>10MB or >4096px),
create low-res previews, split into grid tiles, and generate grid
overlay images with tile labels
- Refactor ImageParser.parse() to integrate large image processing pipeline
- Enable SVG-to-PNG conversion in utils.py (cairosvg/wand)
- Rename ImageConfig.max_dimension to preview_max_dimension and add new
config fields: max_file_size_mb, max_tile_size_mb, max_tile_dimension_px,
tile_overlap_px, large_image_threshold_dimension
- Update ov.conf.example with new image config options
* fix(parse): correct tile dimension comment from 1024px to 2048px
* fix(parse): fix tile label path in grid overlay to include tiles/ directory
* fix(parse): register missing image extensions for ImageParser
TIFF, ICO, DIB, ICNS, SGI, JP2 were not in IMAGE_EXTENSIONS, causing
them to fallback to TextParser. All are supported by PIL.
* fix(parse): preserve PNG format for tiles instead of always converting to JPEG
* fix(parse): address review feedback for large image processing
- Wire config.image to ImageParser in ParserRegistry (was missing)
- Remove unnecessary preview creation for small images (broke LA mode PNG)
- Enforce max_tile_size_mb on tiles with quality reduction and resize fallback
- Remove 64-tile hard cap that conflicted with max_tile_dimension_px
- Add comment explaining why original file is not saved for large images
* refactor(parse): remove max_tile_size_mb as it is a soft suggestion
max_tile_size_mb was a soft constraint that was not enforced
consistently. Remove it from config, constants, and all enforcement
logic. Tile dimension (max_tile_dimension_px) remains the sole constraint.
* fix(parse): use CJK-capable font for grid overlay labels
The old font loading only tried macOS-specific paths and fell back to
PIL's default bitmap font, which cannot render CJK characters in
filenames. Add a cross-platform CJK font lookup that covers Linux
(Noto/Droid/WQY/DejaVu), macOS (PingFang), and Windows (MSYH/SimSun).
* fix(parse): convert non-VLM-supported image formats to PNG on save
Image formats like TIFF, ICO, DIB, ICNS, SGI, JP2 are not recognized
by VLM backends (OpenAI/LiteLLM/VolcEngine only support PNG/JPEG/GIF/
WebP/BMP) or by embedding_utils for image vectorization. When a file
with one of these extensions is parsed, convert it to PNG and use a
.png extension so that downstream pipelines see consistent data.
SVG files (already PNG-converted via cairosvg) also get the .png
extension for the same reason.
* fix(parse): import io for SVG conversion
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
---------
Co-authored-by: openviking <openviking@example.com>
Add FeishuParser that supports importing Feishu cloud documents
(docx, wiki, sheets, bitable) into the knowledge base via URL.
Supports:
- Docx documents via Blocks API with attribute-driven block detection
- Wiki pages (auto-resolve to underlying document type)
- Spreadsheets via Sheets API
- Bitable (multi-dimensional tables) via Bitable API
- Embedded sheet views inside docx documents
- Generic text extraction fallback for unknown block types
Design:
- Uses lark-oapi SDK for all API calls (auth, pagination, etc.)
- Attribute-driven block detection: inspects which SDK attribute is
populated rather than hard-coding 50+ block type integer constants
- Follows convert-then-parse pattern: Feishu -> Markdown -> MarkdownParser
- Lazy imports to avoid breaking when lark-oapi is not installed
- FeishuConfig for credentials (env vars or ov.conf)
Integration:
- URL routing in media_processor.py for feishu.cn/larksuite.com
- Parser registration with ImportError guard
- parse-feishu optional dependency group in pyproject.toml
Tested end-to-end: Feishu URL -> parse -> VikingFS -> L0/L1 generation
-> vectorization -> semantic search, all working.
* feat(parse): add support for legacy .doc and .xls file formats
Add LegacyDocParser using olefile to extract text from Word 97-2003
binary .doc files via OLE2 stream parsing with piece table support
and multi-level fallbacks.
Extend ExcelParser to handle .xls files using xlrd, branching the
parse logic based on file extension while keeping openpyxl for
.xlsx/.xlsm.
New dependencies: olefile>=0.47, xlrd>=2.0.1
* fix(parse): handle date and boolean cell types in xlrd .xls parsing
Check cell.ctype for XL_CELL_DATE and XL_CELL_BOOLEAN to avoid
outputting raw float serial numbers for dates and numeric 0/1 for
booleans.
* fix(parse): harden legacy .doc/.xls parsers per code review
excel.py:
- Enable formatting_info=True so xlrd detects date cells properly
- Add on_demand=True and release_resources() for memory efficiency
- Handle all xlrd cell types: DATE (with time), BOOLEAN, ERROR, BLANK, EMPTY
- Display integers without trailing .0
- Extract cell formatting to _format_xls_cell static method
legacy_doc.py:
- Add 50MB stream size cap to prevent DoS from crafted files
- Cap ccpText at 10M chars to prevent memory exhaustion
- Add FIB version check (require Word 97+ / nFib >= 0x00C1)
- Add minimum buffer length check before struct.unpack_from
- Fix Grpprl skip loop to prevent spin on zero-length entries
- Add _clean_word_text for \x0B (soft break) and \x0C (section break)
- Log warnings for pieces extending beyond stream bounds
- Cap fallback extract to max stream size
* fix: make rust CLI (ov) commands match python CLI (openviking) exactly - add top-level wait/status/health commands
* fix: ov cli plays same as py cli (ls, tree)
* Refactor media parsers to subdirectory structure with validation
* Enhance CLI robustness: validate add-resource path exists and detect unquoted spaces
* Fix unescaped spaces in paths by replacing \ with space
* Sanitize URI components to replace spaces and special chars with underscores
* feat: auto organize audio and image and video files
* Update media parsers to use original filenames and folder names with extensions
* Optimize MediaParser section for readability
* feat: vlm optimization for image
* feat: vlm optimization for image
* feat: vlm optimization for image
* refactor: move media content understanding to SemanticProcessor
- Add parse/parsers/media/utils.py with media helpers
- Refactor ImageParser.parse(), AudioParser.parse(), VideoParser.parse() to remove content understanding, keep only metadata extraction
- Update SemanticProcessor._generate_single_file_summary() to handle media types and call media utils for summary generation
- Update TreeBuilder._get_base_uri() to use media utils
- Update ResourceNode.get_abstract() and get_overview() to check meta for abstract/overview
- Add debug logs and error handling
* refactor: split _generate_single_file_summary to add _generate_text_summary
- Add _generate_text_summary function for text file processing
- Update media utils functions to accept llm_sem and use it to limit concurrent calls
- Update _generate_single_file_summary to call _generate_text_summary and media utils functions
- Fix import ordering
- Fix issue where _generate_file_summaries was creating a new semaphore, now each _generate_single_file_summary handles its own
* feat: vlm optimization for image
* feat: vlm optimization for image
---------
Co-authored-by: openviking <openviking@example.com>
* feat: add directory parsing support to OpenViking
- Implemented DirectoryParser to handle local directories with mixed document types.
- Enhanced add_resource function to support directory imports with options for including, excluding, and ignoring specific directories.
- Updated client and service layers to forward additional parsing options.
- Added unit tests for DirectoryParser to ensure correct functionality and error handling.
- Improved user feedback with rich table summaries for processed, failed, unsupported, and skipped files during directory imports.
* docs: update README.md to include directory import instructions for add.py
* style: reformat files to pass CI code formatting
* style: reformat files to pass CI code formatting
* feat: add markitdown-inspired file parsers
Add support for parsing additional file formats inspired by microsoft/markitdown:
- Word (.docx) - using python-docx
- PowerPoint (.pptx) - using python-pptx
- Excel (.xlsx) - using openpyxl
- Audio (.mp3, .wav, .m4a, etc.) - metadata extraction using mutagen
- EPub (.epub) - using ebooklib
- ZIP (.zip) - iterate contents
All parsers convert content to markdown and delegate to MarkdownParser
for tree structure creation, following OpenViking's parser pattern.
Dependencies added to pyproject.toml:
- python-docx, python-pptx, openpyxl
- ebooklib, beautifulsoup4
- mutagen
Includes comprehensive tests for all new parsers.
Refs: markitdown-parsers
* feat: make markitdown parsers built-in capabilities
Move parser dependencies from optional to main dependencies.
Register parsers directly without graceful fallback.
Remove optional registration infrastructure.
Parsers now built-in:
- Word (.docx) via python-docx
- PowerPoint (.pptx) via python-pptx
- Excel (.xlsx) via openpyxl
- EPub (.epub) via ebooklib
- ZIP (.zip) via built-in zipfile
- Audio (.mp3, .wav, etc.) via mutagen
* refactor: align parsers with existing ecosystem
Set source_format on ParseResult like TextParser does:
- word: source_format = 'docx'
- powerpoint: source_format = 'pptx'
- excel: source_format = 'xlsx'
- epub: source_format = 'epub'
- zip: source_format = 'zip'
- audio: source_format = 'audio'
All parsers now follow the same pattern as existing TextParser
and PDFParser for consistency.
* fix: align markitdown parsers with ecosystem patterns
Critical fixes:
- Remove duplicate zip_archive.py (conflicting ZipParser class name)
- Use zip_parser.py as canonical ZIP parser (follows TextParser pattern)
- Fix parse_content() to delegate to MarkdownParser instead of raising
ValueError (all parse_content tests were broken)
- Set parser_name on all ParseResult outputs (was missing)
- Set source_format AFTER MarkdownParser call (was being overwritten)
- Accept ParserConfig in all parser __init__ (ecosystem consistency)
- Add .xlsm to ExcelParser supported_extensions
- Fix AudioParser._format_size to match ZipParser format (500.0 B)
- Fix pyproject.toml urllib3 indentation corruption
- Add tests/parse/conftest.py with VikingFS test fixture
- Rewrite tests to actually pass and cover registry integration
All 23 tests passing.
* fix: improve markitdown parser consistency and add real-file tests
Critical fixes:
- WordParser: preserve table position in document order (was appending
all tables at end, losing context). Walk document body XML in order
instead of iterating paragraphs then tables separately.
- PowerPointParser: replace magic number (type == 1) with proper
PP_PLACEHOLDER enum constants, also handle CENTER_TITLE.
- AudioParser: add Vorbis/FLAC/OGG tag extraction (previously only
handled ID3 and MP4 formats). Tries all format mappings with dedup.
- ZipParser: replace emoji in tree view with plain text markers
for robustness in text processing pipelines.
- TextParser: set parser_name='TextParser' on parse_content results
for consistency with all other parsers.
- __init__.py: export all new parser classes for public API.
Tests (16 new, 39 total):
- Real .docx/.xlsx/.pptx file creation and parsing
- EPub HTML-to-markdown conversion edge cases
- ZIP bad-file error handling and no-emoji tree view
- AudioParser Vorbis tag extraction and edge cases
- WordParser can_parse() extension matching
* feat: update for rebase and remove audio redundant
* feat: rollback unexpected change
* chore: remove redundant mutagen for audio file