Commit Graph
376 Commits
Author SHA1 Message Date
bot-of-qin-ctxandbot-of-qin-ctx 5de59a1141 fix(tasks): keep parent semantic refresh asynchronous (#4041)
Co-authored-by: bot-of-qin-ctx <222410531+qin-ptr@users.noreply.github.com>
2026-08-17 00:39:18 +08:00
baojun-zhang 84c0895c44 fix(queuefs): skip add-resource lock replay after persisted result (#4007) 2026-08-14 17:12:38 +08:00
Qin Haojie 46f0d60c60 fix(pack): restore account backups without deleting target-only data (#4003) 2026-08-14 16:26:32 +08:00
Jiahui Zhou 62fbf68e84 Discard invalid write-time search tags (#4000) 2026-08-14 14:25:01 +08:00
dingbenandTRAE CLI 3cd1d4e9ac fix(vikingdb): normalize all date_time range filters in API key client (#3973)
* fix(vikingdb): normalize all date_time range filters in API key client

OpenViking compiles TimeRange down to the internal `range` DSL, but the
commercial VikingDB data plane (Bearer API-key auth) expects `time_range`
for date_time fields and `range` only for numeric fields. The API-key
client does not run the local engine's filter conversion, so `range`
nodes on date_time fields were sent verbatim and mis-handled.

Normalize `range` -> `time_range` for every schema date_time field by
reusing the canonical VALID_TIME_FIELDS constant, covering both
`created_at` and `updated_at` instead of hardcoding a single field name.
Numeric `range` nodes and nested boolean filter structure are preserved,
and filters already emitted as `time_range` pass through unchanged. Only
the request body `filter` is rewritten; upsert/update data is untouched.

Add regression tests covering the converted created_at/updated_at date
filters, an unchanged numeric filter, and time_range idempotency.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(vikingdb): normalize date_time filters in AK/SK client

The API-key client already rewrites `range` filter nodes on date_time
fields to VikingDB's `time_range` operator, but the AK/SK-signed
`VolcengineCollection` shares the same commercial data-plane endpoints
and had the identical latent bug: `TimeRange` expressions compile down
to the internal `range` DSL, which the commercial API only accepts for
numeric fields.

Mirror the API-key fix in `VolcengineCollection._data_post` so both
auth modes normalize `range` -> `time_range` for `created_at`/`updated_at`
while leaving numeric `range` nodes untouched. Add AK/SK coverage for
both date_time fields and for idempotency of already-`time_range` input.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-14 11:09:46 +08:00
chenjwandqin-ctx ed1bd4b897 refactor(memory): 统一 V3 提取并提升会话提交与评测稳定性 (#3346)
* fix(memory): disable unsupported tool and skill extraction

* refactor(memory): retire SessionCompressorV2

* docs: design service import cycle fix

* fix(import): break QueueFS service import cycle

* update

* docs: design memory overview lock coverage fix

* fix(memory): cover overview files in update leases

* docs: design session commit default concurrency 50

* perf(queue): raise session commit concurrency to 50

* docs: revise session commit concurrency design

* docs: plan session commit default 8

* perf(queue): default session commit concurrency to 8

* fix(bot): disable cron during eval chat

* docs: design memory link lock stabilization

* docs: plan memory link lock stabilization

* fix(memory): stabilize link update lock coverage

* docs: cover remapped post-group link locks

* docs: design plain-content patch validation

* docs: design first failing patch diagnostics

* fix: report actual failing patch block

* fix(memory): remap replacement links before locking

* fix(bot): include trusted identity in health probe

* test: consolidate memory contract coverage

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-13 21:48:49 +08:00
Qin Haojie 9d5646169b fix(observer): keep empty retrievals diagnostic-only (#3985)
Remove the cumulative retrieval error state and preserve zero-result metrics
without using retrieval yield to determine component health.
2026-08-13 21:46:31 +08:00
Qin Haojie ec18c1dcd8 fix(observer): treat empty retrievals as healthy (#3975)
Record actual search execution errors at the service boundary and use those
errors, rather than empty-result rates, to determine retrieval health.
2026-08-13 21:19:36 +08:00
Jiajie - He/him/his 3577f77423 fix(compile): salvage partial output on timeout and iteration limits (#3948)
* fix(service): break startup circular imports with lazy exports

* fix(fs): avoid root semantic refresh when removing resource scope

* fix(compile): preserve existing wiki links

* fix(sdk): extend HTTP timeout for blocking batch writes

* fix(compile): salvage workspace output on runtime timeout

* feat(compile): support runtime timeout and salvage partial output

* fix: expand compile task and output limits

* revert file

* fix(compile): harden salvage and deadline handling

* update

* fix(compile): address salvage review feedback

* fix(compile): normalize escaped salvage links
2026-08-12 18:51:35 +08:00
Qin Haojie c96fbcb85f fix(session): 避免后序归档阻塞队列 Worker (#3944)
* fix(session): 避免归档任务阻塞队列 Worker

后序归档不再占用 Worker 等待前序任务,并根据 QueueFS work 识别和跳过无法恢复的孤儿归档。

* fix(session): 仅调度队首归档任务

同一 Session 只将最早的未完成归档放入 QueueFS,后续归档在前序结束后再依次入队,并兼容升级前已入队任务。

* fix(session): 降低队首归档调度的存储读取

用 QueueFS 运行时索引判断 Session 是否已有归档任务,正常完成后直接调度相邻 Archive,避免每次 Commit 和任务结束都扫描完整历史目录。

* fix(session): 恢复每个归档任务独立入队
2026-08-12 17:35:46 +08:00
MaojiaSheng b84395d6dd refactor: vikingfs.py (#3947)
* refactor(vikingfs): 将 4513 行的 openviking/storage/viking_fs.py 单体文件拆分为一个包,包含 8 个 mixin 子模块。同时将 _sync_topdown_recursive 的 diff+mv/rm 逻辑从 semantic_processor.py 提取到新的 VikingFS.sync_tree 方法中。SyncDiff 替代了旧的 DiffResult。

* refactor(vikingfs): 将 4513 行的 openviking/storage/viking_fs.py 单体文件拆分为一个包,包含 8 个 mixin 子模块。同时将 _sync_topdown_recursive 的 diff+mv/rm 逻辑从 semantic_processor.py 提取到新的 VikingFS.sync_tree 方法中。SyncDiff 替代了旧的 DiffResult。
2026-08-12 12:35:18 +08:00
Zayn JarvisandClaude Opus 5 4920297ccc feat(mcp): add write/edit/tree tools for viking:// as agent working directory (#3936)
* fix(storage): keep non-memory appends free of memory trailers

ContentWriteCoordinator._write_in_place routed every append through
MemoryFileUtils, which strips the existing trailing newline and appends
a reserved MEMORY_FIELDS metadata trailer, even for resource/skill files
where MEMORY_FIELDS is not a reserved format (see content_visibility).
Append to non-memory files now concatenates raw content instead, matching
POSIX append semantics and the documented visibility rules.

* feat(mcp): add write tool with exact-string edit support

Agents could not use viking:// as a working directory through MCP: no
tool could create or update file content. Add a write tool covering full
writes (mode=replace as create-or-overwrite, append, strict create) and
targeted edits (a list of {old_string, new_string, replace_all}
exact-string replacements applied in order, all-or-nothing), following
the Write/Edit conventions of common agent harnesses.

Edits read via read_visible and write back through the content-write
coordinator, so memory metadata trailers are preserved and semantic /
vector re-indexing triggers as with any other write. Parent directories
are created automatically by the storage layer. Descriptions spell out
writable scopes (resources, user memories/resources, agent) and the
wait=true knob for read-after-write search consistency.

Also update the stale tool-count comment in app.py and the MCP tool
tables in the en/zh guides (13 -> 14 tools).

* feat(mcp): add tree tool, split targeted edits into edit tool

tree renders the recursive directory tree under a viking:// URI,
indented by depth with file sizes, for whole-layout orientation;
level_limit/node_limit bound the output and include_abstract adds
per-file summaries. Missing directories report "(nothing under ...)"
instead of an error, matching the read tool's convention.

edit(uri, old_string, new_string, replace_all) takes over the targeted
exact-string replacement that previously lived in write's edits array,
matching the classic Edit tool signature harnesses already train on.
write now only does full-content writes (content + mode), removing the
mutually-exclusive content/edits schema ambiguity. Edits still read via
read_visible and write back through the content-write coordinator, so
memory metadata trailers are preserved and re-indexing triggers as with
any other write.

* test(plugin): update canonical MCP tool list for tree/write/edit

The marketplace test pins the server-registered MCP tool list; add the
new tree, write, and edit tools to fix plugin-tests CI.

* feat(storage): support plain files at the user scope root

Agents treating viking:// as a working directory naturally drop files
like viking://user/zeus-persona.md at the user root, but the write
coordinator only accepted the memories/ and resources/ subtrees.

Two changes make that work:

- Namespace shorthand: a dotted first segment under viking://user/ is a
  file name, not a user id (canonical user ids are dot-free by
  convention), so viking://user/zeus-persona.md now canonicalizes to
  viking://user/<current-user>/zeus-persona.md, matching how the
  reserved memories/resources/skills segments already shorthand.
  Dot-free segments still address an explicit user, and an exact match
  with the current user id still wins.

- Coordinator: plain files directly under the user root (or in
  non-managed subdirectories) anchor their semantic refresh at the
  parent directory. The managed subtrees skills/, peers/, privacy/ and
  sessions/ remain read-only with an actionable error message.

* fix(namespace): narrow user-root shorthand to text-file extensions

Review on #3936 (codex /review-pr) flagged that treating any dotted
segment as a user-root file shorthand would silently re-route canonical
URIs for valid dotted user ids (e.g. alice.smith) into the current
user space. Shorthand now triggers only when the first segment ends
in a common text-file extension; dotted or email-style user ids keep
resolving as canonical user ids. Adds regression tests pinning both
behaviors.

* fix(mcp): resolve user URIs against current user

* test(mcp): pin plain-file writes directly at the user root

The user-root shorthand exists so an agent can drop viking://user/persona.md
into its workspace, but every new test went through an intermediate directory
(viking://user/project/zeus-persona.md), leaving the no-directory shape — the
one that anchors the write coordinator's refresh at the user root itself —
uncovered. Add the missing case.

Also correct the write tool docstring: the create-extension allowlist applies
to any newly created file, including one created by mode="replace" falling
back to create, not only to an explicit mode="create".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:11:27 +08:00
Qin Haojie 31e01c58a2 feat(admin): 清理已删除用户数据 (#3924)
* feat(admin): 清理已删除用户数据

删除用户时立即撤销身份,并通过持久队列完成用户数据清理。

* fix(admin): 删除用户时清理任务记录

* fix(admin): 避免过早判定用户任务取消失败
2026-08-11 11:18:40 +08:00
Qin Haojieandsponge225 b877ababa5 perf(queue): 流式调度语义向量化任务 (#3636)
* perf(queue): stream semantic vectorization tasks

* fix(queue): isolate semantic work context

* feat(config): make parse concurrency configurable

* test(queue): update semantic vectorization fakes

* fix(queue): correct semantic vectorization conflict resolution

---------

Co-authored-by: sponge225 <1670519171@qq.com>
2026-08-10 21:24:11 +08:00
Qin Haojie 7abd6ab249 refactor(client): remove Python embedded mode (#3712)
* refactor(client): remove Python embedded mode

Consolidate Python consumers on the HTTP SDK while keeping shared server and storage capabilities unchanged.

* refactor(client): remove obsolete embedded leftovers
2026-08-10 18:00:00 +08:00
zgyandqin-ctx 75a1447dc2 perf(queuefs): defer embedding content materialization (#3871)
* Avoid queueing full content for local vector backends

* perf(queue): defer full content materialization

* perf(queuefs): keep deferred content payload empty

* perf(queuefs): separate embedding input from full-text content

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-10 17:06:11 +08:00
Oneflyandqin-ctx 10fd775ac1 fix(storage): exclude zero-byte files from index expectations (#3904)
* fix(storage): exclude zero-byte files from index expectations

* test(storage): remove index consistency tests

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-10 16:01:52 +08:00
7f6085a2f9 feat(memory): support event tag filtering (#3850)
* feat(memory): support event tag filtering

Add session-level default event tags, commit-time overrides, durable queue propagation, and first-write vector index tagging. Include config update APIs and coverage for serialization, concurrency, extraction, and HTTP behavior.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* feat(memory): expose event tags in SDKs and CLI

Add session default tag configuration, config updates, and commit-time event tag overrides across embedded Python, standalone Python, TypeScript, Go, and the Rust CLI. Preserve explicit empty-tag semantics and document each public interface.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(sdk): align legacy session tag APIs

Forward commit-time event tags through the legacy Python HTTP shims and align BaseClient session signatures without adding a new abstract-method requirement for existing subclasses.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* feat(session): allow updating auto-commit policy

Extend PATCH session config to atomically update event tags and auto-commit settings. Merge policy objects by field, use explicit null to disable automatic commits, preserve omitted fields, and expose the contract across SDKs and CLI.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(session): align session config interfaces

Replace the generic session create config JSON flag with explicit event-tag and auto-commit options. Preserve omitted, object, and null auto-commit semantics across HTTP, embedded clients, SDKs, and CLI, reject ambiguous null policy fields, and handle nullable event configuration consistently.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* test(session): trim redundant event tag tests

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-10 11:58:02 +08:00
bianbiandashen 60a02927be fix(vectordb): fill all missing schema fields in fix_fields_data (#3674)
fix_fields_data backfills schema fields absent from a row's data with their
defaults, but it first short-circuited on
`len(field_data_dict) >= len(field_meta_dict)`, using field count as a proxy
for "all schema fields are present".

That proxy is wrong: a row can have as many keys as the schema (or more) while
still missing a specific field — e.g. a row written before a new field was
added that also carries an extra internal/non-schema key. In that case the
fill loop was skipped and the missing field was silently omitted rather than
defaulted, so downstream reads see an incomplete record.

Remove the count guard. The loop already skips fields that are present, so
complete inputs are returned unchanged; only genuinely missing fields are now
filled.

Add regression tests: a missing field with matching key count is backfilled
(default value and type default), and complete data is returned unchanged.
2026-08-08 01:34:11 +08:00
bianbiandashen 0c16a4e478 fix(vectordb): accept WGS-84 boundary coordinates in parse_geo_point (#3671)
The range checks used strict inequalities (< instead of <=), rejecting
the valid boundary values defined by the WGS-84 geographic standard:
  - latitude  ±90  (the North and South Poles)
  - longitude ±180 (the antimeridian / International Date Line)

Any resource whose geo_point is exactly on a pole or the date line cannot
be indexed, and a geo_range query centred on those boundaries raises
ValueError instead of executing the search.

Change both checks to <=. Values strictly outside the valid interval
(e.g. 181, -91) continue to raise ValueError as before.

Add a regression test covering all four boundary endpoints and all four
out-of-range coordinates.
2026-08-07 20:50:02 +08:00
Qiaochu Hu 0205914dc2 fix(storage): propagate VikingFS.mkdir backend errors instead of swallowing them (#3731)
The except block in VikingFS.mkdir() had no re-raise, so any backend
failure that was not an already-exists error (permission denied, quota
exceeded, I/O errors, lock-lease violations) — and even already-exists
errors with exist_ok=False — was silently discarded and mkdir() returned
as if the directory had been created. Callers on the write hot path
(ovpack import, parsers, session, privacy) then write into a directory
that may not exist, and the original actionable error is lost.

Re-raise the original exception unless it is an already-exists error
tolerated by exist_ok=True.

Also update tests/misc/test_mkdir.py, which still mocked fs.agfs.mkdir
even though mkdir() now goes through the AsyncAGFSClient wrapper
(self._async_agfs) — the swallowed-attribute-error made the stale tests
pass/fail for the wrong reasons. Add regression tests covering error
propagation for both exist_ok values.
2026-08-07 20:49:57 +08:00
Qiaochu Hu ecab57e1cc fix(storage): enumerate all entries when mv copies a directory (#3732)
_copy_dir_through_vikingfs() drives the copy phase of mv() for non-temp
directories, but enumerated the source with the agent-facing ls() default
node_limit=1000. Any directory level with more than 1000 visible entries
was copied only partially, and mv() then unconditionally deleted the
source recursively — permanently losing every entry past the cap, while
the vector index (remapped via the uncapped _collect_uris) kept pointing
at URIs that no longer exist anywhere.

Pass the module's LS_ALL_NODES sentinel, which exists precisely for
internal callers that must enumerate an entire directory.

Add a regression test that fails without the fix.
2026-08-07 20:48:55 +08:00
bianbiandashen ef577033b5 fix(vectordb): enforce the UINT16 length contract for list<string> elements (#3681)
The row serializer writes each string's byte length as a UINT16 prefix. The
scalar `string` field guards this contract — a >65535-byte value raises a
clean, field-attributed ValueError. The `list<string>` element path uses the
identical UINT16 prefix but had no such guard, so an oversized element instead
raised a raw `struct.error: 'H' format requires 0 <= number <= 65535` from
deep inside struct.pack_into — a different exception type, naming no field.

Callers that catch ValueError (matching the documented scalar contract) do not
catch this, and the error gives no clue which field/element overflowed. Apply
the same bounds check to list<string> elements so inclusion in a list does not
silently downgrade the type-checked contract the scalar path upholds.

Add a regression test asserting the clean ValueError for an oversized element
and that an in-bounds (incl. multibyte) list still round-trips.
2026-08-07 20:00:22 +08:00
Jiahui Zhou 1d02a72b2b Remove qdrant and opengauss vector backends (#3872) 2026-08-07 19:57:45 +08:00
Jiahui Zhou 6f43a4040c feat: support processing mode for content write (#3615) 2026-08-05 21:42:33 +08:00
Jiahui Zhou d2056e971d Feat/session auto commit v2 (#3736) 2026-08-05 16:18:49 +08:00
Kchen 8d1d52fe5d 资源导入:支持解析后不拆分文档 (#3645) 2026-08-05 11:34:10 +08:00
agent 2ced3f1539 feat(agent-evolution): track experience trajectory lineage (#3727)
* feat(agent-evolution): track experience trajectory lineage

* fix(agent-evolution): return matched trajectory records

* fix(agent-evolution): stabilize lineage pagination
2026-08-04 16:15:06 +08:00
Kchenandchenpengfei c7ee8e753a feat(queue): make ExternalParse worker concurrency configurable (#3710)
* feat(config): make external parse concurrency configurable

* refactor(config): move external parse concurrency to queue workers

---------

Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
2026-08-04 12:10:01 +08:00
Jiahui ZhouandTRAE CLI f857349d6b feat(vectordb): add text field type for large strings (#3725)
The bytes_row STRING type uses a uint16 length prefix, capping a single
field at 65535 bytes. Add a new TEXT field type (enum value 9) that mirrors
STRING semantics (utf-8 str round-trip) but uses a uint32 length prefix,
lifting the per-field limit to ~4GB.

TEXT is added only at the physical bytes_row layer, across all serializers
that must stay byte-identical: the C++ engine (bytes_row.h/.cpp), the abi3
boundary (abi3_engine_backend.cpp, decoding to str not bytes), the pure
Python fallback (store/bytes_row.py), and the engine API (_python_api.py).
Existing types and the CandidateData.fields field are untouched, so old
on-disk data stays readable without reindex.

Fields opt into the new type via metadata={"field_type": FieldType.text}.

Add TestTextFieldType covering >65535-byte round-trips, py<->cpp cross
read/write, binary consistency, and metadata-based declaration.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-04 11:33:18 +08:00
agent 1494dc3052 feat(agent-evolution): improve provenance, history, and account settings (#3695)
* feat(agent-evolution): record trajectory mapping in snapshot commits

* feat(agent-evolution): apply global switch immediately

* fix(snapshot): hide memory fields in visible history

* fix(agent-evolution): preserve batch snapshot provenance

* feat(agent-evolution): scope runtime settings by account

* fix(agent-evolution): serialize experience snapshot finalization
2026-08-03 20:52:46 +08:00
Jiahui ZhouandTRAE CLI 5ab434528e refactor(vectordb): unify random recall via client-generated vector (#3706)
Route the query() random-sampling branch through search_by_vector with a
client-generated random vector (config.embedding.dimension) so every
backend behaves consistently, instead of each backend's server-side
search_by_random. This also makes Qdrant/OpenGauss truly random rather
than a deterministic scroll/scan.

Drop the raise_on_error path entirely per request: query(), the
Collection wrapper, and HttpCollection.search_by_random no longer take
raise_on_error, and delete() no longer requests it. As a result, HTTP
filter-based deletion id lookups now return empty on non-200 instead of
raising.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-03 20:24:00 +08:00
Haoyu ZhangandQin Haojie 3f3554256b feat: 支持基于火山方舟的音视频多模态理解 (#3563)
* feat: add audio and video understanding via VLM

* docs: design media resource guards

* fix: bound media staging concurrency

* fix: cap unknown-size media staging

* test: stage media in routing fake

* test: exercise media staging callbacks

* test: trim media understanding coverage

* chore: 清理实现计划文档

* fix: 修复多凭证切换问题

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-08-03 16:44:22 +08:00
Hao Zheandzhiheng.liu af914ca27b fix(core): harden privacy and background failure handling (#3548)
* fix(server): stop exporting raw query strings and buffering zip responses in observability

Sweep findings: B-03, B-13. Prevent query secrets from reaching traces and keep ZIP responses streaming.

(cherry picked from commit d8ac3dc33b)

* fix(session): tolerate missing/corrupt archive in Phase-2 replay

(NotFoundError / _ArchiveMessagesCorruptError) on a missing or corrupt
archive messages.jsonl instead of returning []. That PR added skip-on-
missing tolerance to the read path (_get_uncovered_archive_messages) and to
resume_queued_commit, but not to the Phase-2 commit replay path
(_prepare_phase2_archive_messages), which calls _read_archive_messages
unguarded while rolling earlier failed archives into the current commit.

Consequence: a terminally-failed earlier archive whose messages.jsonl is
missing/corrupt (legacy "no messages" terminal data, or produced by #3417's
own archive_read terminal path) makes every subsequent commit's Phase-2
extraction raise -> caught by _run_memory_extraction's except -> the current
archive is terminal-failed too. Because the poisoned archive is only removed
from replay once "covered" (which requires a later archive to complete), and
no later archive can ever complete, the session's memory extraction is
permanently poisoned. Raw messages are safe, but extraction is stuck.

Fix: wrap the replay-loop _read_archive_messages call in the same tolerance
_get_uncovered_archive_messages already uses -- skip + warn on not-found
(_is_storage_not_found) and on _ArchiveMessagesCorruptError, re-raise real
storage failures. The skipped archive stays in covered_failed so the current
archive's .done marks it covered, clearing the poison permanently.

Adds a regression test asserting the replay skips a failed archive with a
missing messages.jsonl (and marks it covered) instead of raising, and that a
real storage failure still propagates.

Follow-up to #3417.

(cherry picked from commit 5b8ec9e68a)

* fix(client): align client surfaces without leaking memory metadata

Reconstructs the client-parity work from upstream PR #3439 on current main and strips reserved memory metadata before line slicing in both embedded and HTTP reads.

Based-on: 48b411d58c

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>

* fix(index): propagate semantic vectorization failures safely

Reconstructs upstream PR #3437 on current main, carries enqueue failures through SemanticDagExecutor, and drains the attempt's embedding tracker before retry-visible failure propagation.

Based-on: 02387deb09

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>

* fix(core): close privacy and embedding failure gaps

* fix(memory): strip repeated metadata trailers

* fix(core): close public memory visibility gaps

* ci: skip embedding-dependent resource test without secrets

---------

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
2026-08-03 16:07:35 +08:00
Qin Haojie 9295a3b955 fix(session): remove actor scope from session lifecycle (#3661)
Keep sessions user-scoped and remove legacy agent fallback that could leak an actor view into commit memory writes.
2026-07-31 19:52:19 +08:00
Qin Haojie 09e42aa739 fix(resource): restore queue status for waited imports (#3658) 2026-07-31 15:40:27 +08:00
zgy c241a2a043 chore: reduce import log noise (#3657) 2026-07-31 15:31:44 +08:00
zgy 49b182045b refactor(parser): Refactor code summaries to fixed skeleton-first routing (#3568)
* Refactor code summary skeleton routing

* Simplify code skeleton routing configuration

* Render C tag skeletons as signatures

* Revert "Render C tag skeletons as signatures"

This reverts commit 8e342055f8.

* Simplify fixed code skeleton summary route

* Inline process skeleton rendering

* Simplify code skeleton routing entrypoints

* Fix code summary review issues

* Address final code summary review feedback

* Route failed tags skeletons to LLM fallback

* Restore CUDA and TS extension routing

* Improve code skeleton query coverage

* Route semantic code detection through skeleton support

* Move process skeleton engine into ast package

* Admit skeleton-supported files during directory scan

* Align code summary docs after main merge

* Reduce code skeleton fallback log verbosity

* chore: require grep-ast 0.9.0
2026-07-31 11:38:57 +08:00
baojun-zhang 7c956f23bc feat(config): add compatible default timeout for ragfs pathlock (#3641)
- add storage.agfs.pathlock.lock_timeout_secs
- use pathlock default timeout instead of hardcoded zero in wrapper
- map legacy storage.transaction.lock_timeout when new config is unset
- remote redolog by using  persistent `session_commit` queue.
2026-07-31 11:23:27 +08:00
Qin Haojie fd42b1ad92 feat(tasks): support task cancellation (#3577)
* feat(tasks): support task cancellation

* refactor(tasks): scope cancellation to current user

* feat(cli): support task cancellation

* refactor(tasks): make cancellation queue-aware

* refactor(tasks): simplify cancellation bookkeeping

* test: remove task cancellation coverage

* refactor(tasks): trim cancellation coordination

* fix(tasks): contain cancellation to owned work

* feat(tasks): persist resource source metadata

* fix(tasks): handle cancelled work consistently

* refactor(tasks): make completion queue-aware

* fix(tasks): persist terminal state before queue ack

* test(tasks): remove added lifecycle tests

* docs(tasks): document task cancellation
2026-07-30 20:34:27 +08:00
zgy 48559ab4f1 Fix temp cleanup for repository tasks directories (#3630) 2026-07-30 19:41:14 +08:00
Eurakaxun 44c6df2622 perf: retrieval, import, LangChain, and session-context optimizations (#3569) 2026-07-30 10:10:11 +08:00
baojun-zhang 2f9451231e refactor(pathlock):using rust implement instead python (#3602)
* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):optimize unit test code

* refactor(pathlock):optimize encryption create func

* refactor(pathlock):avoid releasing handoffed pathlock on enqueue errors

* fix(pathlock): use owned lease capability and handle S3 create-new 409 as conflict

* fix(ragfs): keep original FsContext for multi-write metadata

* fix(pathlock): resolve lease coverage and CAS handling issues

- detect S3 conditional conflicts from structured service errors
- pass transaction leases when deleting skill roots
- let temp cleanup acquire locks for temp paths
- disambiguate cache and pathlock providers in cache tests
- update temp cleanup lease assertions

* fix(ragfs): bypass pathlock for multi-write metadata

* fix(ragfs): revert pathlock fail-fast design

* fix(ragfs):fix(ragfs): use non-blocking fcntl locks for localfs CAS

* fix(ragfs): serialize heartbeat lease refresh with release and report real conflict kind

* fix(ragfs): preserve conflict kind snapshot and drop unused test scaffolding

* fix(ragfs): preserve conflict kind snapshot and drop unused test scaffolding
2026-07-29 19:45:34 +08:00
Jiahui Zhou 34b5a88971 Feat/add resource tags (#3560)
* feat: allow tags during resource import

* feat: support uploaded resource watches with tags

* feat: add resource tag flags to CLI

* fix: reject uploaded resource watches with tags

* docs: untrack add resource tags design draft

* fix: write add_resource tags during ingest

* docs: move add_resource tags docs into resources api

* fix: tighten add_resource tag ingestion semantics

* fix: address add_resource tag review feedback

* fix: merge resource tags at vector upsert

* refactor: carry add_resource tags with ingest options
2026-07-29 15:43:06 +08:00
baojun-zhang 1841dfed81 Revert "refactor(pathlock):using rust implement instead python (#3557)" (#3597)
This reverts commit 6b538db569.
2026-07-29 11:31:41 +08:00
baojun-zhang 6b538db569 refactor(pathlock):using rust implement instead python (#3557)
* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):optimize unit test code

* refactor(pathlock):optimize encryption create func

* refactor(pathlock):avoid releasing handoffed pathlock on enqueue errors

* fix(pathlock): use owned lease capability and handle S3 create-new 409 as conflict

* fix(ragfs): keep original FsContext for multi-write metadata

* fix(pathlock): resolve lease coverage and CAS handling issues

- detect S3 conditional conflicts from structured service errors
- pass transaction leases when deleting skill roots
- let temp cleanup acquire locks for temp paths
- disambiguate cache and pathlock providers in cache tests
- update temp cleanup lease assertions

* fix(ragfs): bypass pathlock for multi-write metadata

* fix(ragfs): revert pathlock fail-fast design
2026-07-29 11:08:42 +08:00
fujiajie666 c91b0d36f2 feat: implement Skill-driven knowledge compilation for ov compile (#3567)
* feat(compile): implement skill-driven ov compile

Require a Skill and run compile tasks through VikingBot AgentLoop with structured wiki bundle rendering and durable task state.

Add OpenViking batch-write and bot proxy APIs, Python SDK and Rust CLI support, shared link and memory helpers, tests, and a One-Page demo.

* fix(compile): refine defaults, links, and failure handling

* fix(compile): normalize skill tools and degrade gracefully

* feat(compile): support skill-defined artifact outputs

* feat(compile): improve artifact reliability and wiki navigation

* feat(compile): support generating and updating skill packages

* fix(compile): validate OKF frontmatter and catalog page types

* feat(compile): rank target catalog and validate updates lazily

* feat(compile): tag generated wiki files for search

* fix(compile): preserve generated skill artifacts in submissions

* fix(compile): enforce fixed toolset and workspace artifact submissions

* fix(skills): preserve nested metadata in skill frontmatter

* fix(compile): normalize wiki paths and citation line breaks

* docs(examples): remove outdated compile demos

* docs(api): document compile and batch-write endpoints

* fix(content): allow arbitrary resource files in batch writes

* fix(compile): harden task lifecycle, auth, and execution

* fix(compile): disable direct exec by default

* fix(compile): allow file-only tasks when exec is disabled

* fix(compile): prevent task lock leaks
2026-07-28 20:33:25 +08:00
Jiahui Zhou 5d1ba45be4 Feat/add resource processing mode (#3566)
* feat: add resource processing mode

* fix: keep semantic artifacts in vectors-only add resource

* test: support processing mode in api test client

* docs: document add resource processing mode

* fix: align processing mode after resource ingestion refactor

* feat: expose processing mode in TypeScript SDK

* fix: preserve add resource compatibility
2026-07-28 20:09:06 +08:00
Rocke Dongandqin-ctx cfddab1a23 fix(storage): preserve source files when rm index cleanup fails (#3519)
* fix(storage): preserve source files when rm index cleanup fails

* test(storage): trim redundant removal cases

* test(storage): remove PR-specific regression tests

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-28 16:03:21 +08:00
agent 8391d3a758 feat: add global Agent Evolution switch and HTTP usage sink (#3223)
* feat: add per-user agent evolution settings

* simplify Agent Evolution user settings

* fix: preserve agent evolution client compatibility

* feat(snapshot): add path diff API

* feat(snapshot): expose path diff in clients and CLI

* fix(agent-evolution): gate case memory production

* fix(agent-evolution): preserve configuration compatibility

* feat(usage): add built-in HTTP sink

* fix(agent-evolution): address PR review findings

* fix(usage): isolate HTTP outbox by destination

* docs(usage): define CountRecord HTTP mapping

* docs(usage): plan CountRecord HTTP implementation

* feat(usage): emit CountRecord over HTTP

* docs(agent-evolution): design global switch

* docs(agent-evolution): plan global switch migration

* feat(agent-evolution): make production switch global

* fix(agent-evolution): preserve embedded defaults

* test(agent-evolution): cover failed archive policy replay

* docs(agent-evolution): clarify embedded compatibility

* docs(agent-evolution): expose global switch in example config

* refactor(agent-evolution): align global setting terminology

* fix(agent-evolution): preserve session skill extraction

* fix(usage-reporter): capitalize count record keys
2026-07-27 20:09:38 +08:00