* fix(task): recover add-resource jobs after restart
Persist asynchronous add-resource work in QueueFS so interrupted jobs can resume instead of leaving tasks running forever.
* fix(queue): omit parser args from prepared jobs
* fix(queue): fail when semantic source is missing
* fix(session): return in-flight archive messages in get_session_context (#3129)
Seed latest_completed_index from 0 instead of commit_count so that
archives whose Phase 1 has completed (messages written, commit_count
advanced) but whose Phase 2 is still running (.done not yet written)
are treated as pending rather than already completed.
The release/0.3.x implementation seeded from 0 and did not have
this bug; the regression was introduced when commit_count was
adopted as the seed value.
* test(session): deterministic regression for pending archive context (#3129)
Replace the monkey-patched commit_async test with a direct
filesystem-state test that sets up the post-Phase-1 archive
(messages.jsonl present, commit_count advanced, no .done marker)
and asserts get_session_context still surfaces the archived
messages.
This is deterministic regardless of the queue-worker architecture
because it creates the archive state directly via the mock AGFS
and loads a fresh session from that state, never calling
commit_async or touching the session compressor.
* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest
Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.
- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit
Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.
* fix(vectordb): reject stale index replacements
* fix(vectordb): harden bulk rebuild lifecycle
---------
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
* feat(parse): add large image processing for image parser
- Add large_image_processor.py: detect large images (>10MB or >4096px),
create low-res previews, split into grid tiles, and generate grid
overlay images with tile labels
- Refactor ImageParser.parse() to integrate large image processing pipeline
- Enable SVG-to-PNG conversion in utils.py (cairosvg/wand)
- Rename ImageConfig.max_dimension to preview_max_dimension and add new
config fields: max_file_size_mb, max_tile_size_mb, max_tile_dimension_px,
tile_overlap_px, large_image_threshold_dimension
- Update ov.conf.example with new image config options
* fix(parse): correct tile dimension comment from 1024px to 2048px
* fix(parse): fix tile label path in grid overlay to include tiles/ directory
* fix(parse): register missing image extensions for ImageParser
TIFF, ICO, DIB, ICNS, SGI, JP2 were not in IMAGE_EXTENSIONS, causing
them to fallback to TextParser. All are supported by PIL.
* fix(parse): preserve PNG format for tiles instead of always converting to JPEG
* fix(parse): address review feedback for large image processing
- Wire config.image to ImageParser in ParserRegistry (was missing)
- Remove unnecessary preview creation for small images (broke LA mode PNG)
- Enforce max_tile_size_mb on tiles with quality reduction and resize fallback
- Remove 64-tile hard cap that conflicted with max_tile_dimension_px
- Add comment explaining why original file is not saved for large images
* refactor(parse): remove max_tile_size_mb as it is a soft suggestion
max_tile_size_mb was a soft constraint that was not enforced
consistently. Remove it from config, constants, and all enforcement
logic. Tile dimension (max_tile_dimension_px) remains the sole constraint.
* fix(parse): use CJK-capable font for grid overlay labels
The old font loading only tried macOS-specific paths and fell back to
PIL's default bitmap font, which cannot render CJK characters in
filenames. Add a cross-platform CJK font lookup that covers Linux
(Noto/Droid/WQY/DejaVu), macOS (PingFang), and Windows (MSYH/SimSun).
* fix(parse): convert non-VLM-supported image formats to PNG on save
Image formats like TIFF, ICO, DIB, ICNS, SGI, JP2 are not recognized
by VLM backends (OpenAI/LiteLLM/VolcEngine only support PNG/JPEG/GIF/
WebP/BMP) or by embedding_utils for image vectorization. When a file
with one of these extensions is parsed, convert it to PNG and use a
.png extension so that downstream pipelines see consistent data.
SVG files (already PNG-converted via cairosvg) also get the .png
extension for the same reason.
* fix(parse): import io for SVG conversion
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
LiteLLMRerankClient.rerank_batch wrapped each document as {"text": d} and read
result items via getattr(item, ...). That works for Cohere-style object results
but breaks Voyage through litellm: Voyage's rerank API rejects dict-wrapped
documents (400: 'documents' is not a valid string), and litellm returns Voyage
results as plain dicts, so getattr(item, "index") misses and rerank silently
falls back to a no-op.
- Pass documents as plain strings (litellm.rerank expects List[str]).
- Add _result_field() to read index/relevance_score from dict- or object-shaped
result items.
Verified end-to-end against voyage/rerank-2.5 (scores now applied). Adds
tests/unit/models/rerank/test_litellm_rerank.py covering the plain-string
documents contract and both dict- and object-shaped results.
Co-authored-by: Michael Tarleton <mtarleton@istation.com>
FastMCP derives tool input schemas from Python type hints, so Optional/
Union parameters become anyOf nodes with no top-level type, and nested
models become $ref/$defs. Valid JSON Schema, but clients that forward
these schemas verbatim to strict function-calling APIs break: Gemini's
OpenAPI 3.0 subset rejects the whole request (400: schema didn't specify
the schema type field), and n8n's JSON-schema-to-Zod conversion can
silently fall back and drop tool arguments entirely.
Rewrite the advertised schemas after registration: drop null branches,
collapse unions to their most general branch, inline $refs, and ensure
every node carries an explicit type. Runtime argument validation still
uses the original function signatures, so union parameters keep
accepting every branch (e.g. read still takes a bare URI string even
though the schema advertises an array).
Also type recall's other_peer_penalty honestly as
Optional[Union[float, Dict[str, float]]] instead of Optional[Any],
which produced an empty {} schema node.
* Batch add-message writes across memory plugins
* fix(plugins): stop batch enqueue at first failure and bump plugin versions
Keep pending-queue entries a contiguous prefix when a mid-stream enqueue
fails: consumers mark the first sent+queued payloads as captured, so a
queued entry after a gap could silently drop the gapped message.
Bump claude-code 0.4.3, codex 0.7.3, opencode 0.2.3, cursor/trae 0.1.2 so
existing installs pick up the batch write path.
* fix(openclaw): restore peer_role routing
The plugin refactor stopped applying peer_role to message attribution and data-plane actor headers, causing assistant mode to tag user messages and route every request as the agent. Restore role-aware routing across context and tool flows.
* fix(openclaw): close peer routing isolation gaps
* fix(openclaw): separate message peers from actor scope
* feat(connector): delegate add_resource imports to external Connector
Opt-in integration that routes add_resource data fetching and parsing
to external Connector service; the Connector stages source data and
calls back into OV through the standard add_resource pipeline.
- add ConnectorClient wrapping the control plane's inner doc/add and
task/info endpoints
- add [connector] config section: enable, connector/tracker endpoint
URLs, timeout_seconds, poll_interval_ms, allowed_add_types
- route add_resource via Connector when enabled and args.add_type is
in allowed_add_types; otherwise fall back to the standard pipeline
with an info log
- track imports as connector_import TaskRecords and poll Connector
task status in the background until terminal state or timeout
* fix(connector): delegate add_resource imports to external Connector