* fix(task): recover add-resource jobs after restart
Persist asynchronous add-resource work in QueueFS so interrupted jobs can resume instead of leaving tasks running forever.
* fix(queue): omit parser args from prepared jobs
* fix(queue): fail when semantic source is missing
* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest
Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.
- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit
Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.
* fix(vectordb): reject stale index replacements
* fix(vectordb): harden bulk rebuild lifecycle
---------
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
* feat(storage): optimize glob func
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* fix(localfs): offload blocking fs operations to spawn_blocking
* feat(glob): cap glob api default node_limit at 256
* feat(sdk): add node_limit options for glob in python and go SDKs
The embedding handler classified 401/403 auth errors as transient and
re-enqueued them, tripping the circuit breaker and cycling the message
forever. add-resource holds the resource tree lock and blocks --wait
until embeddings complete, so a permanently-failing credential (the dummy
keys on no-secrets fork-PR CI) hung the lock and the request
indefinitely; every add-resource retry then hit "Resource is busy".
Route ERROR_CLASS_AUTH to terminal failure (mark failed, no re-enqueue,
no breaker trip) so the embedding tracker drains, the request completes,
and the tree lock releases. The resource is left un-vectorized; a reindex
recovers it once credentials are valid.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
[EN]
viking_fs.ls() defaults to node_limit=1000 to keep agent-facing tool
output from flooding the model's context. Several internal system
operations on the ingest / summary / vectorize path call ls() to
enumerate a directory's children and inherited that 1000 cap, so
importing a directory with more than 1000 entries silently processed
only the first 1000 and dropped the rest.
Observed: a 6,221-document import produced exactly 1000 subdirectories.
The namespace already held an earlier import, so temp->final
materialization took the incremental sync path
(_sync_topdown_recursive -> list_children) instead of the atomic
whole-directory move; list_children's ls() truncated the 6,221 temp
children to 1000 and the remaining 5,221 were dropped when temp was
deleted.
Fix: introduce a shared LS_ALL_NODES sentinel in viking_fs and pass it
explicitly at every internal call site that must observe every child.
ls()'s default stays 1000, so agent-facing listings are unchanged.
Call sites fixed:
- DirectoryParser._merge_temp / _recursive_move (parser temp merge)
- SemanticProcessor._sync_topdown_recursive (temp->final sync)
- SemanticProcessor._process_memory_directory (memory dirs)
- SemanticDagExecutor._list_dir (summary DAG dispatch + recursion)
- Summarizer.list_top_children (semantic-unit enqueue)
- embedding_utils.index_resource (per-directory file indexing)
Tests: tests/storage/test_ingest_ls_node_limit.py reproduces the >1000
truncation for both the temp->final sync materialization and the summary
DAG enumeration (red before, green after). Existing fakes updated to
accept the node_limit kwarg production now passes.
[中文]
viking_fs.ls() 默认 node_limit=1000,用于避免 agent 工具输出刷爆模型上下文。
但入库 / 摘要 / 向量化链路上多处内部系统调用 ls() 枚举目录子节点时也继承了
这个上限,导致目录条目超过 1000 时只处理前 1000 个、其余被静默丢弃。
现象:6221 篇文档入库后,目标命名空间下只剩正好 1000 个子目录。由于该命名
空间已存在更早的入库结果,temp->final 物化走了增量同步路径
(_sync_topdown_recursive -> list_children)而非原子整目录搬移;list_children
的 ls() 把 6221 个 temp 子节点截断到 1000,其余 5221 个在 temp 清理时丢失。
修复:在 viking_fs 中引入共享哨兵 LS_ALL_NODES,在每一处必须枚举全部子节点的
内部调用显式传入;ls() 默认值仍为 1000,agent-facing 的列目录行为不变。修复的
调用点见上方 Call sites。若不一并修摘要 DAG / sync,即使 raw 物化修好,1001 篇
之后的 L0/L1 摘要与向量索引仍会卡在 1000。
测试:tests/storage/test_ingest_ls_node_limit.py 复现 temp->final 物化同步与摘要
DAG 枚举两处的 >1000 截断(修复前 red、修复后 green);现有 fake 已更新以接受生产
代码新传入的 node_limit 参数。
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(storage): optimize tree func
* feat(storage): add test case && fix tree show parent dir problem
* feat(storage): format code
* feat(storage): fix check problem
* feat(storage): remove redundant unit test
* feat(storage): fix code review issue
* feat(storage): fix code review issue
Cache async SDK clients per event loop to avoid cross-loop reuse in worker threads.
Move memory vectorization into semantic queue refresh and preserve target sync state for resource updates.
* auto-commit before eval 20260509_181850
* auto-commit before eval 20260509_192618
* update
* auto-commit before eval 20260510_005109
* auto-commit before eval 20260510_011832
* auto-commit before eval 20260510_014114
* auto-commit before eval 20260510_022835
* auto-commit before eval 20260510_025048
* auto-commit before eval 20260510_031034
* auto-commit before eval 20260510_143728
* auto-commit before eval 20260510_172705
* auto-commit before eval 20260510_220133
* auto-commit before eval 20260511_115905
* auto-commit before eval 20260511_121959
* auto-commit before eval 20260511_132120
* auto-commit before eval 20260511_161430
* auto-commit before eval 20260511_163606
* auto-commit before eval 20260511_173943
* auto-commit before eval 20260511_175657
* auto-commit before eval 20260511_224347
* auto-commit before eval 20260511_233109
* auto-commit before eval 20260512_104710
* auto-commit before eval 20260512_111256
* auto-commit before eval 20260512_181905
* auto-commit before eval 20260512_191540
* auto-commit before eval 20260512_192540
* auto-commit before eval 20260512_195710
* auto-commit before eval 20260513_000746
* auto-commit before eval 20260513_004221
* auto-commit before eval 20260513_004656
* refactor: migrate logger calls to tracer in extract_loop modules
Replace logger.warning/error/info with tracer.error/info in extract_loop
related modules for better observability (console + OpenTelemetry spans).
Modules updated:
- agent_experience_context_provider.py (5 replacements)
- extract_loop.py (4 replacements)
- memory_updater.py (9 replacements)
- session_extract_context_provider.py (4 replacements)
- utils/json_parser.py (7 replacements)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* auto-commit before eval 20260513_123007
* auto-commit before eval 20260513_125305
* auto-commit before eval 20260513_135421
* auto-commit before eval 20260513_141013
* auto-commit before eval 20260513_143455
* auto-commit before eval 20260513_145401
* auto-commit before eval 20260513_163345
* auto-commit before eval 20260514_105906
* auto-commit before eval 20260514_112912
* auto-commit before eval 20260514_120308
* auto-commit before eval 20260514_122022
* auto-commit before eval 20260514_134800
* auto-commit before eval 20260514_135615
* auto-commit before eval 20260514_135818
* auto-commit before eval 20260514_142941
* auto-commit before eval 20260514_162401
* auto-commit before eval 20260514_231859
* auto-commit before eval 20260515_104122
* auto-commit before eval 20260515_122140
* auto-commit before eval 20260515_122942
* auto-commit before eval 20260515_144941
* auto-commit before eval 20260515_154736
* auto-commit before eval 20260515_181643
* auto-commit before eval 20260515_182727
* auto-commit before eval 20260515_183056
* auto-commit before eval 20260515_183652
* auto-commit before eval 20260515_183825
* auto-commit before eval 20260515_202731
* auto-commit before eval 20260516_001144
* auto-commit before eval 20260516_011749
* auto-commit before eval 20260516_015903
* auto-commit before eval 20260516_020505
* auto-commit before eval 20260516_130701
* auto-commit before eval 20260516_144342
* auto-commit before eval 20260516_151043
* Harden memory graph rendering and patch guidance.
Escape embedded graph data for script safety, add a vis-network load guard, tighten graph layout defaults, and clarify SEARCH guidance so patch content stays bound to the target file/page context.
🤖 Generated with [Aiden x Claude Code]
Co-Authored-By: Aiden
* auto-commit before eval 20260517_005258
* auto-commit before eval 20260517_012903
* auto-commit before eval 20260517_014036
* auto-commit before eval 20260517_015726
* auto-commit before eval 20260517_024952
* auto-commit before eval 20260517_032518
* auto-commit before eval 20260517_135114
* auto-commit before eval 20260517_143238
* auto-commit before eval 20260517_154858
* auto-commit before eval 20260517_200556
* auto-commit before eval 20260517_215025
* fix: keep memory storage plain and render graph links on display
Store memory bodies as plain text in VikingFS and move link rendering to graph display so repeated writes no longer persist nested markdown links. Also tighten link renderer path handling so cross-user relative paths are rejected and strip_links preserves viking and absolute targets.
🤖 Generated with [Aiden x Claude Code]
Co-Authored-By: Aiden
* auto-commit before eval 20260518_001945
* fix: invert selected graph node colors
Make the currently selected memory node use a light background with dark text so it stands out against the dark graph theme.
🤖 Generated with [Aiden x Claude Code]
Co-Authored-By: Aiden
* auto-commit before eval 20260518_011327
* update
* auto-commit before eval 20260518_161813
* auto-commit before eval 20260518_165104
* auto-commit before eval 20260518_174259
* update
* auto-commit before eval 20260518_224834
* auto-commit before eval 20260518_233319
* auto-commit before eval 20260518_235712
* auto-commit before eval 20260519_135952
* fix memory patch failure logging
Keep dry-run patch validation from emitting a misleading patch_handler warning, and record skipped field updates from MemoryUpdater where the failure is handled.
🤖 Generated with [Aiden x Claude Code]
Co-Authored-By: Aiden
* auto-commit before eval 20260519_213142
* 语言修正
* fix(memory): fan out links for shared page ids
Expand _resolve_links so shared page ids resolve across every operation URI instead of collapsing to a single path. Align the page-id and extract-loop tests with the current API contract and the multi-URI link behavior.
🤖 Generated with [Aiden x Claude Code]
Co-Authored-By: Aiden
* auto-commit before eval 20260520_195141
* update language
* 语言修正,日语相关
* 语言修正,针对时区,小语言,兼容windows系统时区
* auto-commit before eval 20260520_215911
* auto-commit before eval 20260520_222335
* update
* style(memory): clean up formatter drift
Apply the remaining formatter-driven cleanup in the memory modules so the working tree stays clean before the next behavior changes. This keeps helper signatures and string literals aligned with current lint output.
🤖 Generated with [Aiden x Claude Code]
Co-Authored-By: Aiden
* 日语修正
* 更新注释
* update
---------
Co-authored-by: chenjunwen <chenjunwen@bytedance.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Introduce typed lock leases for semantic queue handoff so resource, memory, and reindex flows can share or transfer lock ownership without releasing caller-owned locks prematurely.
* fix(semantic): ensure memory processing always reports completion status
_process_memory_directory() had early return paths that could bypass
report_success()/report_error() in on_dequeue(), leaving the queue's
in_progress counter permanently stuck. This caused the semantic queue
to appear stalled with pending items never being processed.
All code paths now properly propagate to the completion callbacks.
Fixes#864.
* fix(semantic): classify filesystem errors as permanent to prevent infinite retry
Address review feedback: filesystem errors (FileNotFoundError,
PermissionError, IsADirectoryError, NotADirectoryError) are now
classified as permanent by classify_api_error(), so they hit
report_error() instead of being infinitely re-enqueued.
Tests updated to exercise real classifier behavior without mocking.
* test: fix set_callbacks signature for DequeueHandlerBase
DequeueHandlerBase.set_callbacks now takes (on_success, on_requeue, on_error);
the original PR #951 test harness called it with only (on_success, on_error).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(semantic): remove dead _mark_failed helper, add transient-error test
_mark_failed's two call sites were removed when _process_memory_directory
started raising on error. The closure itself was left behind. Delete it —
telemetry failure is now reported by on_dequeue's exception handler via
get_request_wait_tracker().mark_semantic_failed().
Add test_memory_ls_transient_error_requeues to cover the transient branch
of the memory path: a 500-class error from ls() must route through
_reenqueue_semantic_msg() and fire report_requeue() + report_success(),
not report_error(). The previous tests only exercised permanent errors.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: deepakdevp <deepakdevp@gmail.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>