Commit Graph
799 Commits
Author SHA1 Message Date
Zonas ZhouandClaude 6e77291265 feat(pdf): refactor MinerU parsing to the official file_parse API (#3953)
* feat(pdf): refactor MinerU parsing to the official file_parse API

* feat(pdf): remove mineru_api_key from configuration and examples

* feat(pdf): preflight MinerU /health during service initialization

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-17 13:48:49 +08:00
Kchenandchenpengfei eeff5a4973 fix(memory): preserve owner for assistant-only event ranges (#4001)
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
2026-08-14 20:16:09 +08:00
baojun-zhang 84c0895c44 fix(queuefs): skip add-resource lock replay after persisted result (#4007) 2026-08-14 17:12:38 +08:00
Qin Haojie 46f0d60c60 fix(pack): restore account backups without deleting target-only data (#4003) 2026-08-14 16:26:32 +08:00
Jiahui Zhou 4d140d6482 Discard invalid write-time search tags (#4005) 2026-08-14 15:22:58 +08:00
996128abcc fix(session): split JSONL on newline only, not Unicode line boundaries (#3984) (#3988)
* fix(session): split JSONL on newline only, not Unicode line boundaries (#3984)

* test(session): consolidate unicode JSONL regression coverage

---------

Co-authored-by: mac <bishopapril850965@yahoo.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-14 14:46:43 +08:00
Jiahui Zhou 62fbf68e84 Discard invalid write-time search tags (#4000) 2026-08-14 14:25:01 +08:00
dingbenandTRAE CLI 3cd1d4e9ac fix(vikingdb): normalize all date_time range filters in API key client (#3973)
* fix(vikingdb): normalize all date_time range filters in API key client

OpenViking compiles TimeRange down to the internal `range` DSL, but the
commercial VikingDB data plane (Bearer API-key auth) expects `time_range`
for date_time fields and `range` only for numeric fields. The API-key
client does not run the local engine's filter conversion, so `range`
nodes on date_time fields were sent verbatim and mis-handled.

Normalize `range` -> `time_range` for every schema date_time field by
reusing the canonical VALID_TIME_FIELDS constant, covering both
`created_at` and `updated_at` instead of hardcoding a single field name.
Numeric `range` nodes and nested boolean filter structure are preserved,
and filters already emitted as `time_range` pass through unchanged. Only
the request body `filter` is rewritten; upsert/update data is untouched.

Add regression tests covering the converted created_at/updated_at date
filters, an unchanged numeric filter, and time_range idempotency.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(vikingdb): normalize date_time filters in AK/SK client

The API-key client already rewrites `range` filter nodes on date_time
fields to VikingDB's `time_range` operator, but the AK/SK-signed
`VolcengineCollection` shares the same commercial data-plane endpoints
and had the identical latent bug: `TimeRange` expressions compile down
to the internal `range` DSL, which the commercial API only accepts for
numeric fields.

Mirror the API-key fix in `VolcengineCollection._data_post` so both
auth modes normalize `range` -> `time_range` for `created_at`/`updated_at`
while leaving numeric `range` nodes untouched. Add AK/SK coverage for
both date_time fields and for idempotency of already-`time_range` input.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-14 11:09:46 +08:00
chenjwandqin-ctx ed1bd4b897 refactor(memory): 统一 V3 提取并提升会话提交与评测稳定性 (#3346)
* fix(memory): disable unsupported tool and skill extraction

* refactor(memory): retire SessionCompressorV2

* docs: design service import cycle fix

* fix(import): break QueueFS service import cycle

* update

* docs: design memory overview lock coverage fix

* fix(memory): cover overview files in update leases

* docs: design session commit default concurrency 50

* perf(queue): raise session commit concurrency to 50

* docs: revise session commit concurrency design

* docs: plan session commit default 8

* perf(queue): default session commit concurrency to 8

* fix(bot): disable cron during eval chat

* docs: design memory link lock stabilization

* docs: plan memory link lock stabilization

* fix(memory): stabilize link update lock coverage

* docs: cover remapped post-group link locks

* docs: design plain-content patch validation

* docs: design first failing patch diagnostics

* fix: report actual failing patch block

* fix(memory): remap replacement links before locking

* fix(bot): include trusted identity in health probe

* test: consolidate memory contract coverage

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-13 21:48:49 +08:00
Guoxuterandguoxuter 5d2ffc1d4f fix: align v7 query planner prompt contract (#3983)
Co-authored-by: guoxuter <j7azwflq4h@gmail.com>
2026-08-13 21:47:29 +08:00
Qin Haojie 9d5646169b fix(observer): keep empty retrievals diagnostic-only (#3985)
Remove the cumulative retrieval error state and preserve zero-result metrics
without using retrieval yield to determine component health.
2026-08-13 21:46:31 +08:00
Qin Haojie ec18c1dcd8 fix(observer): treat empty retrievals as healthy (#3975)
Record actual search execution errors at the service boundary and use those
errors, rather than empty-result rates, to determine retrieval health.
2026-08-13 21:19:36 +08:00
Jiahui ZhouandTRAE CLI 482434ef6e feat(reindex): support tag updates (#3964)
* feat(reindex): support tag updates

Add replace and append tag modes to reindex vector rebuilds, preserve omission-aware behavior, and propagate options through background tasks and namespace rebuilds. Align Python, TypeScript, Go, and CLI interfaces with tests and documentation.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(reindex): lock file targets exactly

Use an exact path lock for existing file targets while retaining tree locks for directories and prune-orphans scopes. Add a regression test for single-file reindex.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(reindex): handle prune file targets

Use exact locks for existing file targets in prune-orphans mode while retaining tree scope for missing targets. Document the existing Go ReindexOptions wait semantics and add lock regression coverage.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-13 13:41:12 +08:00
MaojiaShengandTRAE CLI 5aed7f72b4 fix(media): downsample oversized image model inputs (#3965)
* fix(embedding): downsample oversized image inputs

Keep imported image resources unchanged while avoiding provider-side multimodal embedding failures for oversized images. The embedding path now builds a temporary downsampled image data URI when image bytes exceed the shared large-image limits.

Move reusable image size thresholds into media_limits so both parser-side large image handling and embedding-side input preparation depend on a common utility instead of embedding_utils importing parser internals.

Add vectorize_file coverage confirming large image embedding inputs are resized and the stored resource bytes are preserved.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(media): downsample oversized image model inputs

Keep imported image resources unchanged while avoiding provider-side multimodal failures for oversized images. Shared image input preparation now builds temporary downsampled bytes for model requests when image bytes exceed the configured large-image limits.

Apply the model-input downsampling to both semantic image summary generation and embedding image data URI construction, so directory and code repository imports can preserve original images while sending provider-compatible inputs.

Move reusable image size thresholds into media_limits so parser-side large image handling, VLM summary generation, and embedding preparation share common limits without embedding_utils importing parser internals.

Always convert downsampled model images to RGB before JPEG encoding so Pillow-openable modes such as LA or I;16 do not fall back to the original oversized bytes.

Add coverage confirming VLM image summaries, vectorize_file embedding inputs, and JPEG-incompatible image modes are resized while stored resource bytes are preserved.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-12 22:24:57 +08:00
Jiajie - He/him/his 3577f77423 fix(compile): salvage partial output on timeout and iteration limits (#3948)
* fix(service): break startup circular imports with lazy exports

* fix(fs): avoid root semantic refresh when removing resource scope

* fix(compile): preserve existing wiki links

* fix(sdk): extend HTTP timeout for blocking batch writes

* fix(compile): salvage workspace output on runtime timeout

* feat(compile): support runtime timeout and salvage partial output

* fix: expand compile task and output limits

* revert file

* fix(compile): harden salvage and deadline handling

* update

* fix(compile): address salvage review feedback

* fix(compile): normalize escaped salvage links
2026-08-12 18:51:35 +08:00
Qin Haojie c96fbcb85f fix(session): 避免后序归档阻塞队列 Worker (#3944)
* fix(session): 避免归档任务阻塞队列 Worker

后序归档不再占用 Worker 等待前序任务,并根据 QueueFS work 识别和跳过无法恢复的孤儿归档。

* fix(session): 仅调度队首归档任务

同一 Session 只将最早的未完成归档放入 QueueFS,后续归档在前序结束后再依次入队,并兼容升级前已入队任务。

* fix(session): 降低队首归档调度的存储读取

用 QueueFS 运行时索引判断 Session 是否已有归档任务,正常完成后直接调度相邻 Archive,避免每次 Commit 和任务结束都扫描完整历史目录。

* fix(session): 恢复每个归档任务独立入队
2026-08-12 17:35:46 +08:00
c1345a1f7e feat(feishu):Support Feishu Drive folder and file imports (#3937)
* Support Lark drive folder and file URLs

* Handle partial Feishu folder import failures

* Fix Feishu drive folder path names

* Tighten Feishu drive folder tests

* test(feishu): consolidate drive import coverage

---------

Co-authored-by: haoxingjun <haoxingjun@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-12 16:01:00 +08:00
Jiahui Zhou c1cc592aaf fix(memory): use last value for duplicate tag keys (#3950) 2026-08-12 15:07:07 +08:00
MaojiaSheng b84395d6dd refactor: vikingfs.py (#3947)
* refactor(vikingfs): 将 4513 行的 openviking/storage/viking_fs.py 单体文件拆分为一个包,包含 8 个 mixin 子模块。同时将 _sync_topdown_recursive 的 diff+mv/rm 逻辑从 semantic_processor.py 提取到新的 VikingFS.sync_tree 方法中。SyncDiff 替代了旧的 DiffResult。

* refactor(vikingfs): 将 4513 行的 openviking/storage/viking_fs.py 单体文件拆分为一个包,包含 8 个 mixin 子模块。同时将 _sync_topdown_recursive 的 diff+mv/rm 逻辑从 semantic_processor.py 提取到新的 VikingFS.sync_tree 方法中。SyncDiff 替代了旧的 DiffResult。
2026-08-12 12:35:18 +08:00
Zayn JarvisandClaude Opus 5 4920297ccc feat(mcp): add write/edit/tree tools for viking:// as agent working directory (#3936)
* fix(storage): keep non-memory appends free of memory trailers

ContentWriteCoordinator._write_in_place routed every append through
MemoryFileUtils, which strips the existing trailing newline and appends
a reserved MEMORY_FIELDS metadata trailer, even for resource/skill files
where MEMORY_FIELDS is not a reserved format (see content_visibility).
Append to non-memory files now concatenates raw content instead, matching
POSIX append semantics and the documented visibility rules.

* feat(mcp): add write tool with exact-string edit support

Agents could not use viking:// as a working directory through MCP: no
tool could create or update file content. Add a write tool covering full
writes (mode=replace as create-or-overwrite, append, strict create) and
targeted edits (a list of {old_string, new_string, replace_all}
exact-string replacements applied in order, all-or-nothing), following
the Write/Edit conventions of common agent harnesses.

Edits read via read_visible and write back through the content-write
coordinator, so memory metadata trailers are preserved and semantic /
vector re-indexing triggers as with any other write. Parent directories
are created automatically by the storage layer. Descriptions spell out
writable scopes (resources, user memories/resources, agent) and the
wait=true knob for read-after-write search consistency.

Also update the stale tool-count comment in app.py and the MCP tool
tables in the en/zh guides (13 -> 14 tools).

* feat(mcp): add tree tool, split targeted edits into edit tool

tree renders the recursive directory tree under a viking:// URI,
indented by depth with file sizes, for whole-layout orientation;
level_limit/node_limit bound the output and include_abstract adds
per-file summaries. Missing directories report "(nothing under ...)"
instead of an error, matching the read tool's convention.

edit(uri, old_string, new_string, replace_all) takes over the targeted
exact-string replacement that previously lived in write's edits array,
matching the classic Edit tool signature harnesses already train on.
write now only does full-content writes (content + mode), removing the
mutually-exclusive content/edits schema ambiguity. Edits still read via
read_visible and write back through the content-write coordinator, so
memory metadata trailers are preserved and re-indexing triggers as with
any other write.

* test(plugin): update canonical MCP tool list for tree/write/edit

The marketplace test pins the server-registered MCP tool list; add the
new tree, write, and edit tools to fix plugin-tests CI.

* feat(storage): support plain files at the user scope root

Agents treating viking:// as a working directory naturally drop files
like viking://user/zeus-persona.md at the user root, but the write
coordinator only accepted the memories/ and resources/ subtrees.

Two changes make that work:

- Namespace shorthand: a dotted first segment under viking://user/ is a
  file name, not a user id (canonical user ids are dot-free by
  convention), so viking://user/zeus-persona.md now canonicalizes to
  viking://user/<current-user>/zeus-persona.md, matching how the
  reserved memories/resources/skills segments already shorthand.
  Dot-free segments still address an explicit user, and an exact match
  with the current user id still wins.

- Coordinator: plain files directly under the user root (or in
  non-managed subdirectories) anchor their semantic refresh at the
  parent directory. The managed subtrees skills/, peers/, privacy/ and
  sessions/ remain read-only with an actionable error message.

* fix(namespace): narrow user-root shorthand to text-file extensions

Review on #3936 (codex /review-pr) flagged that treating any dotted
segment as a user-root file shorthand would silently re-route canonical
URIs for valid dotted user ids (e.g. alice.smith) into the current
user space. Shorthand now triggers only when the first segment ends
in a common text-file extension; dotted or email-style user ids keep
resolving as canonical user ids. Adds regression tests pinning both
behaviors.

* fix(mcp): resolve user URIs against current user

* test(mcp): pin plain-file writes directly at the user root

The user-root shorthand exists so an agent can drop viking://user/persona.md
into its workspace, but every new test went through an intermediate directory
(viking://user/project/zeus-persona.md), leaving the no-directory shape — the
one that anchors the write coordinator's refresh at the user root itself —
uncovered. Add the missing case.

Also correct the write tool docstring: the create-extension allowlist applies
to any newly created file, including one created by mode="replace" falling
back to create, not only to an explicit mode="create".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:11:27 +08:00
agent 00f3738edb feat(usage): emit resource-scoped experience usage records (#3921)
* feat(usage): expand experience tracking and log schema

* fix(usage): preserve experience count event names

* refactor(agent-evolution): use generic OpenViking tools

* fix(usage): capture generic OpenViking tool events

* feat(skills): guide cross-agent experience retrieval

* fix(usage): address generic tool migration review
2026-08-11 22:20:18 +08:00
Qin Haojie 0ea10190cc fix(memory): keep patch work off event loop (#3934)
Run fuzzy patch application in a worker thread and make merge operations
awaitable so session commits cannot starve request handling.
2026-08-11 20:03:51 +08:00
Zayn Jarvis dcca29364c fix(parse): stop silently dropping markdown YAML frontmatter (#3929)
Frontmatter was parsed into ParseResult.meta and removed from the body, but
that metadata is never persisted, so every ingested markdown file lost its
frontmatter with no way to read the fields back.

Parse frontmatter into meta unconditionally (it still drives doc_title) and
only remove it from the stored body when explicitly configured; that removal
is now off by default.
2026-08-11 14:14:53 +08:00
zihengli cd55ec89a6 feat(assets): support fixed Git commits, explicit targets, and private repository auth (#3703)
* feat/openviking_assets_support_git_commit_id

* feat/openviking_assets_support_to

* feat/private_git_support_watch

* fix: doc_and_ut

* fix: doc_and_ut

* fix: adapt git token url
2026-08-11 14:07:21 +08:00
Qin Haojie 31e01c58a2 feat(admin): 清理已删除用户数据 (#3924)
* feat(admin): 清理已删除用户数据

删除用户时立即撤销身份,并通过持久队列完成用户数据清理。

* fix(admin): 删除用户时清理任务记录

* fix(admin): 避免过早判定用户任务取消失败
2026-08-11 11:18:40 +08:00
Qin Haojieandsponge225 b877ababa5 perf(queue): 流式调度语义向量化任务 (#3636)
* perf(queue): stream semantic vectorization tasks

* fix(queue): isolate semantic work context

* feat(config): make parse concurrency configurable

* test(queue): update semantic vectorization fakes

* fix(queue): correct semantic vectorization conflict resolution

---------

Co-authored-by: sponge225 <1670519171@qq.com>
2026-08-10 21:24:11 +08:00
Qin Haojie 9d5cbae70e fix(examples): align quick start with server mode (#3923) 2026-08-10 20:59:35 +08:00
Qin Haojie 7abd6ab249 refactor(client): remove Python embedded mode (#3712)
* refactor(client): remove Python embedded mode

Consolidate Python consumers on the HTTP SDK while keeping shared server and storage capabilities unchanged.

* refactor(client): remove obsolete embedded leftovers
2026-08-10 18:00:00 +08:00
agent a779c62a0f feat(agent-evolution): add trajectory date filters (#3915)
* feat(agent-evolution): add trajectory date filters

* refactor(agent-evolution): simplify date validation
2026-08-10 17:53:41 +08:00
Jiajie - He/him/his 6a252eba18 fix(session): validate archive names when loading history (#3922) 2026-08-10 17:48:49 +08:00
zgyandqin-ctx 75a1447dc2 perf(queuefs): defer embedding content materialization (#3871)
* Avoid queueing full content for local vector backends

* perf(queue): defer full content materialization

* perf(queuefs): keep deferred content payload empty

* perf(queuefs): separate embedding input from full-text content

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-10 17:06:11 +08:00
8acaf7f872 fix(watch): 避免资源移动期间持有全局锁 (#3833)
* fix(watch): avoid global lock during resource moves

Allow unrelated watch operations to continue while viking_fs.mv is running, while serializing overlapping resource paths with a move fence. Preserve transaction integrity across persistence failures, rollback, and caller cancellation.

* feat(resource): add URI mutation coordinator

* refactor(watch): separate target rewrites from resource moves

* refactor(fs): own resource move watch transaction

* feat(watch): coordinate refreshes with URI mutations

* refactor(core): share URI mutation coordinator

* test(watch): focus URI move coverage

* test(watch): reuse existing URI move coverage

---------

Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-10 16:39:37 +08:00
Yu ZhangandClaude Sonnet 4.6 6617a92cdd fix(task-tracker): delete expired persisted tasks (#3913)
Keep expired terminal tasks cached when persistent deletion fails so cleanup can retry without resurrecting stale records after restart.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-10 16:30:56 +08:00
Jiajie - He/him/his cfd74413cb fix(semantic): reduce unsupported entity hallucinations in summaries (#3908)
* fix(semantic): reduce hallucination in summary/overview prompts

* fix(semantic): harden entity faithfulness in summary prompts
2026-08-10 14:13:09 +08:00
agent bca5a67388 feat(agent-evolution): add trajectory task query (#3856) 2026-08-10 14:05:10 +08:00
t0sakiandTRAE CLI 3160d423f9 fix(retrieval): record first-turn session recalls (#3907)
Materialize session-aware context requests through SessionService before
loading the recall ledger, while retaining the messages.jsonl guard for
partially initialized sessions.

Add coverage for first-turn ledger writes, same-session message capture,
materialization failure recovery, and stateless requests with session
features disabled.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-10 13:55:39 +08:00
7f6085a2f9 feat(memory): support event tag filtering (#3850)
* feat(memory): support event tag filtering

Add session-level default event tags, commit-time overrides, durable queue propagation, and first-write vector index tagging. Include config update APIs and coverage for serialization, concurrency, extraction, and HTTP behavior.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* feat(memory): expose event tags in SDKs and CLI

Add session default tag configuration, config updates, and commit-time event tag overrides across embedded Python, standalone Python, TypeScript, Go, and the Rust CLI. Preserve explicit empty-tag semantics and document each public interface.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(sdk): align legacy session tag APIs

Forward commit-time event tags through the legacy Python HTTP shims and align BaseClient session signatures without adding a new abstract-method requirement for existing subclasses.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* feat(session): allow updating auto-commit policy

Extend PATCH session config to atomically update event tags and auto-commit settings. Merge policy objects by field, use explicit null to disable automatic commits, preserve omitted fields, and expose the contract across SDKs and CLI.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(session): align session config interfaces

Replace the generic session create config JSON flag with explicit event-tag and auto-commit options. Preserve omitted, object, and null auto-commit semantics across HTTP, embedded clients, SDKs, and CLI, reject ambiguous null policy fields, and handle nullable event configuration consistently.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* test(session): trim redundant event tag tests

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-10 11:58:02 +08:00
dingbenandTRAE CLI 9097fef478 feat(server): refresh read-replica API key index via store watcher (#3857)
Read replicas load the API key store once at startup and never rewrite
it, so a user registered/rotated/removed on the writer stays invisible
(new key -> "Invalid API Key"; removed key -> still accepted).

Add an optional background watcher that polls the shared key store and
reloads the in-memory index only when it actually changes:

- APIKeyManager.reload(): strictly read-only refresh that rebuilds state
  and swaps it in atomically, never writing or migrating plaintext keys.
- compute_store_signature(): cheap (path, size, modTime) signature over
  accounts.json + every users.json so the watcher skips unchanged polls.
- ApiKeyAuthPlugin starts/stops the watcher behind api_key_watch_enabled
  (default off) with api_key_watch_interval_seconds; AuthPlugin.shutdown()
  is wired into app shutdown to cancel it cleanly.

Add coverage for reload convergence, read-only/no-migrate guarantees,
uninitialized-store tolerance, signature change detection, and watcher
reload/skip/shutdown behavior.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-10 11:28:52 +08:00
bianbiandashen 3087f943a2 fix(markdown): keep force-split chunks within the token budget (#3672)
_smart_split_content documents that it enforces both a token limit
(max_size) and a hard character limit. But when a single paragraph was
oversized by tokens yet under the character limit, the force-split loop
stepped through it by max_chars only, so it emitted a chunk that still
exceeded max_size tokens.

This is reachable with ordinary long-form CJK text: _estimate_token_count
weights CJK at ~0.7 token/char, so a ~5000-char Chinese paragraph is
~3500 tokens (over the 2048 default) while staying under the char limit,
and was returned as a single over-budget chunk.

Bound the force-split step by min(max_chars, max_size / MAX_TOKENS_PER_CHAR),
where MAX_TOKENS_PER_CHAR is the worst-case (CJK) density already used by
_estimate_token_count, now extracted into a shared constant so the two stay
in sync.

Add regression tests for the token budget and content preservation.
2026-08-08 01:54:06 +08:00
lazyayuan cbc3907791 fix(retrieve): store L1 overview in abstract scalar so Rerank sees L1 text (#3714) 2026-08-08 01:40:32 +08:00
bianbiandashen 60a02927be fix(vectordb): fill all missing schema fields in fix_fields_data (#3674)
fix_fields_data backfills schema fields absent from a row's data with their
defaults, but it first short-circuited on
`len(field_data_dict) >= len(field_meta_dict)`, using field count as a proxy
for "all schema fields are present".

That proxy is wrong: a row can have as many keys as the schema (or more) while
still missing a specific field — e.g. a row written before a new field was
added that also carries an extra internal/non-schema key. In that case the
fill loop was skipped and the missing field was silently omitted rather than
defaulted, so downstream reads see an incomplete record.

Remove the count guard. The loop already skips fields that are present, so
complete inputs are returned unchanged; only genuinely missing fields are now
filled.

Add regression tests: a missing field with matching key count is backfilled
(default value and type default), and complete data is returned unchanged.
2026-08-08 01:34:11 +08:00
bianbiandashen 7b8b33e8f8 fix(markdown): match GitHub anchor slugs for headings with punctuation (#3673)
_gh_slug claims to produce "GitHub-style" heading slugs but collapsed runs
of whitespace (`re.sub(r"\s+", "-", s)`). GitHub's reference slugger
(github-slugger) maps each space to its own hyphen (`.replace(/ /g, '-')`)
and does not collapse.

Because punctuation is stripped before spaces are converted, a heading like
"Foo & Bar" leaves two adjacent spaces where "&" was. GitHub renders this as
"foo--bar", but _gh_slug produced "foo-bar". The intra-document link rewriter
(_rewrite_link) compares _gh_slug(heading) against the link fragment, so an
author-written link such as `guide.md#foo--bar` failed to match its heading
and was left unrewritten after the target doc was split into sections.

Replace `\s+` with `\s` so each whitespace character maps to one hyphen,
matching GitHub. Simple single-space headings are unaffected.

Add regression tests covering punctuation headings and ordinary headings.
2026-08-07 20:56:40 +08:00
bianbiandashen 0c16a4e478 fix(vectordb): accept WGS-84 boundary coordinates in parse_geo_point (#3671)
The range checks used strict inequalities (< instead of <=), rejecting
the valid boundary values defined by the WGS-84 geographic standard:
  - latitude  ±90  (the North and South Poles)
  - longitude ±180 (the antimeridian / International Date Line)

Any resource whose geo_point is exactly on a pole or the date line cannot
be indexed, and a geo_range query centred on those boundaries raises
ValueError instead of executing the search.

Change both checks to <=. Values strictly outside the valid interval
(e.g. 181, -91) continue to raise ValueError as before.

Add a regression test covering all four boundary endpoints and all four
out-of-range coordinates.
2026-08-07 20:50:02 +08:00
Qiaochu Hu 0205914dc2 fix(storage): propagate VikingFS.mkdir backend errors instead of swallowing them (#3731)
The except block in VikingFS.mkdir() had no re-raise, so any backend
failure that was not an already-exists error (permission denied, quota
exceeded, I/O errors, lock-lease violations) — and even already-exists
errors with exist_ok=False — was silently discarded and mkdir() returned
as if the directory had been created. Callers on the write hot path
(ovpack import, parsers, session, privacy) then write into a directory
that may not exist, and the original actionable error is lost.

Re-raise the original exception unless it is an already-exists error
tolerated by exist_ok=True.

Also update tests/misc/test_mkdir.py, which still mocked fs.agfs.mkdir
even though mkdir() now goes through the AsyncAGFSClient wrapper
(self._async_agfs) — the swallowed-attribute-error made the stale tests
pass/fail for the wrong reasons. Add regression tests covering error
propagation for both exist_ok values.
2026-08-07 20:49:57 +08:00
Qiaochu Hu ecab57e1cc fix(storage): enumerate all entries when mv copies a directory (#3732)
_copy_dir_through_vikingfs() drives the copy phase of mv() for non-temp
directories, but enumerated the source with the agent-facing ls() default
node_limit=1000. Any directory level with more than 1000 visible entries
was copied only partially, and mv() then unconditionally deleted the
source recursively — permanently losing every entry past the cap, while
the vector index (remapped via the uncapped _collect_uris) kept pointing
at URIs that no longer exist anywhere.

Pass the module's LS_ALL_NODES sentinel, which exists precisely for
internal callers that must enumerate an entire directory.

Add a regression test that fails without the fix.
2026-08-07 20:48:55 +08:00
bianbiandashen ef577033b5 fix(vectordb): enforce the UINT16 length contract for list<string> elements (#3681)
The row serializer writes each string's byte length as a UINT16 prefix. The
scalar `string` field guards this contract — a >65535-byte value raises a
clean, field-attributed ValueError. The `list<string>` element path uses the
identical UINT16 prefix but had no such guard, so an oversized element instead
raised a raw `struct.error: 'H' format requires 0 <= number <= 65535` from
deep inside struct.pack_into — a different exception type, naming no field.

Callers that catch ValueError (matching the documented scalar contract) do not
catch this, and the error gives no clue which field/element overflowed. Apply
the same bounds check to list<string> elements so inclusion in a list does not
silently downgrade the type-checked contract the scalar path upholds.

Add a regression test asserting the clean ValueError for an oversized element
and that an in-bounds (incl. multibyte) list still round-trips.
2026-08-07 20:00:22 +08:00
Jiahui Zhou 1d02a72b2b Remove qdrant and opengauss vector backends (#3872) 2026-08-07 19:57:45 +08:00
750390b0cf fix(server): map HTTP status for error-envelope return sites in stats/debug/sessions (#3722)
Several routers returned Response(status='error', error=ErrorInfo(...))
directly, so FastAPI shipped them with HTTP 200 instead of the mapped
status from ERROR_CODE_TO_HTTP_STATUS. Web-studio's sessions API only
survived because of a code-based fallback; other clients (e.g. a generic
HTTP retry layer) would treat these as success and never surface the
error.

Switch the seven return sites to error_response() so the canonical
mapping drives the HTTP status, and add the three previously-unmapped
codes (INTERNAL_ERROR, NO_VECTOR_DB, INVALID_FILTER) to the map.

Residual follow-up from PR #1764 (ac9f679a).

Co-authored-by: ming <silverchris@foxmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 19:55:02 +08:00
Ray Tien 830aad6ec9 fix #3859: normalize intent reasoning output (#3864) 2026-08-07 16:56:20 +08:00
baojun-zhang bd5cce09c7 feat(pathlock): adjust pathlock config (#3854) 2026-08-07 13:21:30 +08:00