* fix(parse): distinguish mpegts from TypeScript ts
* fix(parse): tighten mpegts ts routing semantics
* fix(semantic): use file name for media summary type
---------
Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* refactor(parse): simplify resource ingestion routing
Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.
* fix(feishu): preserve sheet and bitable imports
Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.
* fix(parse): keep normalized Feishu content internal
Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.
* fix(feishu): parse bitable blocks embedded in sheets
Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.
* fix(feishu): download bitable attachment images
* refactor(parse): remove unused document converter
* refactor(parse): unify Understanding routing
* docs(parse): mark routing classification points
* docs(parse): complete wait routing flow
* fix(parse): preserve Feishu Base URL scope
* refactor(resource): separate ingestion submission from execution
* fix(resource): reject internal ingestion fields at public entry
* feat(connector): delegate add_resource imports to external Connector
Opt-in integration that routes add_resource data fetching and parsing
to external Connector service; the Connector stages source data and
calls back into OV through the standard add_resource pipeline.
- add ConnectorClient wrapping the control plane's inner doc/add and
task/info endpoints
- add [connector] config section: enable, connector/tracker endpoint
URLs, timeout_seconds, poll_interval_ms, allowed_add_types
- route add_resource via Connector when enabled and args.add_type is
in allowed_add_types; otherwise fall back to the standard pipeline
with an info log
- track imports as connector_import TaskRecords and poll Connector
task status in the background until terminal state or timeout
* fix(connector): delegate add_resource imports to external Connector
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
Keep full background add-resource tasks limited to Git repositories so anti-crawler HTTP pages are parsed by the normal importer instead of failing during early source validation.
* feat: add async task tracking for add-resource/add-skill/write operations
- Return task_id when add-resource/add-skill/write called without --wait
- Add 'ov task status <task_id>' and 'ov task list' CLI commands
- Bridge RequestWaitTracker and TaskTracker via background monitor coroutines
- Format TaskRecord timestamps as ISO 8601 in to_dict()
- Always generate telemetry_id (remove 'not wait' condition)
- Extract _create_write_task helper to eliminate code duplication
- Add unregister_wait_telemetry in _monitor_write_queue for consistency
- Update CLI async prompt from 'ov wait' to 'ov task status <task_id>'
- Add 7 new tests covering async task tracking
* refactor: remove task_id from write operations, keep only add-resource/add-skill
Write operations (write/create_file/write_memory) are primarily called
by internal flows like session commit, which have their own task tracking.
Only add-resource and add-skill are user-facing CLI commands that need
task_id for async progress tracking.
* fix: address PR review feedback
1. Fix async failure path leaking request-scoped tracker/telemetry state
- Add monitor_started flag to ensure cleanup when monitor coroutine
hasn't been launched yet
- Finally block now cleans up if wait or not telemetry_id or not monitor_started
2. Fix TaskRecord timestamp API compatibility break
- Keep original created_at/updated_at as float (backward compatible)
- Add new created_at_iso/updated_at_iso fields with ISO 8601 strings
3. Fix ruff format lint failure on resource_service.py
4. Add regression test for async failure cleanup
- test_add_resource_async_failure_cleans_up_tracker verifies no
RequestWaitTracker or telemetry registry state leaks when processor
raises before task/monitor creation
* fix: add missing unregister_wait_telemetry in add_skill finally block
* fix: remove unused imports in test_add_resource_async_failure_cleans_up_tracker
* fix: queue failure shows as completed and business error creates unreachable task
- _monitor_queue_processing: check error_count from build_queue_status,
mark task as failed when queue processing has errors
- add_resource: skip task creation when process_resource returns
status=error, preventing unreachable ghost tasks
- Improve test_add_resource_async_failure_cleans_up_tracker: patch
internal processor instead of add_resource itself to cover finally
cleanup logic
- Add test_add_skill_async_returns_task_id for add_skill coverage
- Add test_add_resource_business_error_no_task regression test
- Add test_monitor_marks_failed_on_queue_error regression test
* test: move add_skill tests to test_session_task_tracking.py
Move add_skill task tracking tests from test_api_resources.py to
test_session_task_tracking.py where other task tracking tests live,
and add sync no-task-id coverage.
* test: update async task queryable assertion to include failed status
Queue errors now correctly mark task as failed (Bug 1 fix), so the
test assertion must accept 'failed' as a valid terminal status.
* refactor: use explicit return result for business error path
* fix: remove duplicate asyncio imports in test_api_resources.py
* fix: avoid task creation on watch conflict
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(security): clean up code scanning and runtime findings
Harden path and logging boundaries, remove noisy cleanup issues,
and keep observability failures from breaking runtime flows.
* fix(security): close werewolf and feishu validation gaps
Block the remaining path traversal bypass in the werewolf demo,
and validate Feishu hosts on the main parse() entry point.
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
---------
Co-authored-by: openviking <openviking@example.com>
* feat(resource): add watch interval support for resource monitoring
implement resource watch functionality that allows automatic monitoring and re-processing of resources at specified intervals. key features include:
- add watch_interval parameter to resource APIs
- create watch scheduler service for task execution
- handle conflict detection for active watch tasks
- provide watch status query capability
- include comprehensive tests and examples
the watch feature enables periodic automatic updates of resources without manual intervention, improving data freshness for frequently changing content
* feat(resources): add watch status tracking and improve resource processing
- Implement get_watch_status API for tracking resource watch status
- Add immediate persistence for first-time resource additions
- Improve file change detection with size comparison
- Refactor watch scheduler with better concurrency control
- Add test coverage for watch status and resource processing
- Remove unused watch manager references and clean up code
* refactor: improve code style and fix minor issues
- Simplify logging by removing redundant data copying
- Fix syntax errors in docstrings and string literals
- Add new fields to EmbeddingMsg class
- Improve line wrapping and formatting
- Update watch task storage URIs to use hidden files
* refactor(embedding_msg): simplify EmbeddingMsg constructor by removing unused fields
Remove media_uri, media_mime_type and id parameters as they are not used in the implementation
* feat(watch): add backup task recovery and simplify permission check
Add test case for recovering tasks from backup storage when primary is missing
Remove require_owner parameter from _check_permission as it's redundant with the existing role-based checks
* feat(resources): add watch_interval support for resource updates
Add watch_interval parameter to enable periodic resource updates. When target is specified, watch_interval > 0 creates/updates a watch task, while <= 0 disables it. Also simplify resource moving logic in ResourceProcessor by using direct mv operation.
* refactor(watch): remove deprecated get_watch_status functionality
remove get_watch_status method and related tests, update examples to use direct task access
update watch manager to use ConflictError for URI conflicts and include original_role in tasks
add validation for watch_interval requiring target URI
* fix(resource_service): validate watch interval before processing resource
Move watch interval validation earlier in the flow to fail fast when 'to' parameter is missing
* refactor(resource_processor): remove redundant temp_uri assignment