Commit Graph
933 Commits
Author SHA1 Message Date
bot-of-qin-ctxandqin-ctx 49afccd608 refactor(acl): model protection state as modes (#4768)
Replace the derived boolean with an extensible mode so restricted inheritance can become a single additional state.

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-09-07 16:39:32 +08:00
agent 230abf054d fix(tags): remove charset and length restrictions (#4761)
* fix(agent-evolution): preserve legacy experience lineage tags

* fix(tags): relax character restrictions while retaining length limits

* fix(tags): remove search tag format and length restrictions

* fix(tags): retain only comma validation for search tags

* fix(tags): limit relaxation to charset and length checks
2026-09-07 14:31:05 +08:00
baojun-zhang 94ff079f51 feat(storage): add pagination capability for ls/tree. support http-api/cli/python-sdk/go-sdk/typescript-sdk/mcp && CLI/MCP support ls sort capability (#4698)
* feat(storage): add pagination capability for ls/tree. support http-api/cli/python-sdk/go-sdk/typescript-sdk/mcp &&  CLI/MCP support ls sort capability

* feat(storage): add pagination capability for ls/tree.

* feat(storage): add go/typescript sdk pagination capability for ls/tree.
2026-09-07 11:50:15 +08:00
Kchenandchenpengfei b73108dabd fix(storage): 恢复 cp/mv 覆盖合并并缩小路径锁范围 (#4718)
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
2026-09-07 10:10:21 +08:00
zihengli f97ce53002 fix: watch task path conflict error (#4719) 2026-09-06 21:06:08 +08:00
Jiahui ZhouandTRAE CLI 2251aa136a perf(storage): skip unnecessary directory counts in internal stat calls (#4703)
Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-09-04 20:12:52 +08:00
Qin Haojie 37db9c834e feat(vectordb): support safe local schema updates (#4700) 2026-09-04 19:17:47 +08:00
94a10313bf fix: watch task status (#4589)
* Revert "ci: persist BuildKit mount caches across runs (#4421)" (#4467)

* ci: build the release docker image once (#4469)

* ci: build the release docker image once, not twice

release.yml carries a full copy of the docker build/manifest jobs from
build-docker-image.yml. Both fire on the same release: the tag push triggers
build-docker-image.yml while the release event triggers release.yml's copy, so
every release builds the same image 4 times (2 arch x 2 workflows, ~16 min
each) and both pipelines race to write the same v0.4.x and latest tags.

Drop the copy. build-docker-image.yml has to exist anyway (main images, manual
dispatch) and emits the identical tag set for a tag ref.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01247pby58vheLBLmHPV3WKQ

* ci: assert the release's docker image actually landed

The tag push and the release event fire in the same second, so the image is
still ~16 min from existing when release.yml starts. Poll the tag's
build-docker-image run, fail on a missing or failed run, then inspect both
registries for the version tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01247pby58vheLBLmHPV3WKQ

* ci: assert latest points at the released tag too

The version tag existing was never the failure mode; latest silently staying
on the previous release was. Compare digests instead of just checking the
version tag is present.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01247pby58vheLBLmHPV3WKQ

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(parser): adapt AnyDoc 0.2 document model (#4509)

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>

* fix: watch task status

* fix: watch task create when add

* fix: ut

* fix: watch task create when add

* fix: ut

* feat: connector support user resources

* test: reproduce Feishu OAuth v3 refresh failure

* test: cover legacy Feishu refresh tokens

* fix: support Feishu refresh token formats

* fix: add logs

* fix: ut

* fix: feishu doc watch

* fix: ai review

* fix: ai review

* test(cli): copy an actual file in cp integration test

* fix: add some log

---------

Co-authored-by: Zayn Jarvis <zaynjarvis@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: bot-of-qin-ctx <qin_haojie@qq.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-09-04 19:04:23 +08:00
zgy 1ee1219ab8 fix: enable explicit openai multimodal embeddings (#4433) 2026-09-04 17:10:32 +08:00
yangxinxin-7andClaude Opus 5 92ccb0f57a feat(search): allow find without a query when a filter is given (#4694)
An exact-match lookup — finding a record by a tag that carries an external
id, say — has no meaningful query text. Callers were forced to invent one,
which made the similarity score noise and left the result at the mercy of
whatever recall the made-up query happened to produce.

find() now accepts an empty query as long as a filter narrows the search.
The result is then fully determined by that filter, so there is nothing to
embed or rank: the request is resolved from the metadata store directly and
`score` stays 0 rather than a fabricated value callers might sort on.

Scoping goes through a new filter_in_tenant, which reuses the very same
_build_scope_filter the vector path uses. Tenant isolation, ACL grants and
target-directory limits therefore stay identical between the two paths — a
hand-built filter here would be one refactor away from silently losing them.

Validation stays at the top of find() so a bad request is still rejected
before any initialization, and an empty query with no filter is still an
error: without either one there is nothing to narrow by, and returning an
arbitrary slice of the store would be worse than failing.

search() is deliberately left alone: it expands intent from session context,
which has no meaning without a query.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 17:02:52 +08:00
Qin Haojie 02e31f2d66 perf(session): reduce commit phase1 latency (#4687)
* perf(session): reduce commit task-store lock overhead

* perf(session): reduce phase1 fs round trips

* chore(session): trim benchmark diagnostics

* perf(session): avoid extending session write lock scope
2026-09-04 16:24:34 +08:00
AirHua c5755f5ae6 fix(vlm): resync rotated Codex credentials (#4540)
* fix(vlm): resync rotated Codex credentials

Signed-off-by: AirHua-byte <78540954+AirHua-byte@users.noreply.github.com>

* fix(vlm): preserve external auth ownership

Signed-off-by: AirHua-byte <78540954+AirHua-byte@users.noreply.github.com>

---------

Signed-off-by: AirHua-byte <78540954+AirHua-byte@users.noreply.github.com>
2026-09-04 15:26:53 +08:00
DuTao 6020c62ccc feat(bot): support multimodal OpenViking resource reads (#4593)
* feat(bot): support multimodal OpenViking reads

* fix(bot): harden multimodal OpenViking reads

* fix(bot): enforce multimodal request budgets

* fix(bot): prefer recent tool media

* fix(vlm): preserve text tool output serialization

* fix(bot): scope image reads to resource owner
2026-09-04 14:52:25 +08:00
Zayn JarvisandClaude Opus 5 a4aa04cfc6 fix(resource): reject a source with no content at ingestion (#4571)
* fix(resource): reject a source with no content at ingestion

Adding a zero-byte file currently succeeds and produces an empty
resource. With the Understanding API disabled -- the default, since
ParserApiConfig.enable is False -- nothing on the internal parse path
looks at the size, so the file is staged, parsed, and indexed as an
empty entry. Directory imports already refuse the same input:
directory_scan.py skips any zero-byte member as an "empty file". A
single-file import should not disagree with that.

Check it in UnifiedResourceProcessor.prepare, the one point every
ingestion path passes through: prepare_durable_source freezes a source
there before the request touches the tree, and process() calls it for
anything not frozen earlier. So local files, uploads, remote downloads,
git and Feishu sources are all covered by one check, at the moment the
bytes are first in hand.

InvalidArgumentError maps to 400 through ERROR_CODE_TO_HTTP_STATUS.
Where ingestion runs asynchronously -- a plain remote URL with the
Understanding API off is queued, and the response has already been
sent -- the same error fails the task instead of the request.

Deliberately narrow:

- Only zero bytes. A one-byte file is still accepted; this is not a
  minimum-size policy.
- Only regular files. A directory has no meaningful size and is skipped,
  so directory and repository imports containing empty files are
  unaffected.
- A stat failure is left to the normal ingestion path to report.
- content/write is untouched. Creating an empty file there is an explicit
  user action, not an ingestion accident.

The error names the caller's own file rather than the temp working copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(resource): clean rejected temporary sources

* fix(resource): preserve queued error codes

* test(resource): focus empty-source regression coverage

* ci: drop dedicated empty-resource regression step

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 14:08:55 +08:00
Onefly 0f58d62a57 fix(embedder): bound OpenAI HTTP connection pools (#4475)
* fix(embedder): bound OpenAI HTTP connection pools

* fix(embedder): validate and close bounded pools
2026-09-03 16:16:26 +08:00
wutongyuonce b759068924 fix(privacy): serialize config mutations under pathlock and propagate storage errors (#4081)
* fix(privacy): serialize config mutations under pathlock and propagate storage errors

- Hold the AGFS tree pathlock around upsert/activate_version/delete read-modify-write
  so concurrent updates no longer all read empty meta and write version 1.
- Only NotFoundError/FileNotFoundError map to empty results; storage, permission,
  and deserialization errors now propagate instead of being swallowed.
- Regression tests: concurrent upserts keep every version; exists() re-raises
  storage errors.

* fix(privacy): wait for config locks and preserve read errors
2026-09-03 15:35:04 +08:00
zgy 63ba98cadb fix(session): tolerate placeholder memory telemetry URIs (#4598) 2026-09-03 14:53:06 +08:00
Kchenandchenpengfei f6d9dec6b6 feat: 增加事务化文件系统复制能力 (#4185)
* refactor: share vector URI rewrite rules

* feat: add strict vector transfer transactions

* feat: add transactional VikingFS copy

* feat: coordinate filesystem copy service

* feat: expose filesystem copy API

* feat: add ov cp command

* test: cover filesystem copy end to end

* test: compose CLI collection hooks

* fix: support vector copy on volcengine backend

* fix: harden copy transaction boundaries

* fix: retain copied entry in parent semantics

* fix: scan local vectors with supported sort key

* fix: refresh move target parent semantics

* fix: report missing copy target directory

* feat: rebuild transfer parent semantics from target summaries

* fix: refresh both parents after move

* docs: document filesystem copy API

* fix: adapt copy flow to canonical URI boundary

* fix: adapt transfer semantics after upstream rebase

* fix: harden copy and move transaction boundaries

* fix: preserve transfer metadata and direct ACLs

* ci: exercise copy with available capabilities

* fix: restore move vectors after ACL refresh failure

---------

Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
2026-09-03 11:32:10 +08:00
YohanesandYohanes 7ab1f96566 fix(session): honor output language for working memory (#4514)
* fix(session): honor output language for working memory

* fix(session): preserve multiline language detection

---------

Co-authored-by: Yohanes <CryoThrust@users.noreply.github.com>
2026-09-03 11:08:17 +08:00
baojun-zhang 35d0b69636 Revert "feat(s3): add aliyun oss vendor support and config (#4564)" (#4603)
This reverts commit 63c25306af.
2026-09-02 21:34:32 +08:00
Jiahui ZhouandTRAE CLI bbc1f9bab0 Feat/cli unify tags flag (#4599)
* refactor(cli): unify tag flag to --tags for add-resource and reindex

Both add-resource and reindex still used the repeatable --tag flag, while
write/ls/tree/grep/glob already use comma-separated --tags. Align them so
every tag-carrying command accepts --tags k=v,k=v consistently.

- add-resource: --tag (repeatable) -> --tags (comma-separated)
- reindex: --tag (repeatable) -> --tags (comma-separated)
- update the two CLI parse tests accordingly

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* feat(tags): constrain search tag charset and length

Explicit k=v search tags previously only required a single '=' and
non-empty sides, so free-form text (spaces, slashes, non-ASCII) and
unbounded length could reach the vector store. Tighten normalize_search_tag
to keep tags as stable identifiers:

- key/value must match ^[a-z0-9][a-z0-9_.-]*$ (after lower-casing)
- key <= 64 chars, value <= 128 chars
- add unit tests for charset and length boundaries
- document the rules in retrieval/content API docs

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-09-02 20:02:13 +08:00
spud 85b4923d06 fix(server): improve invalid config diagnostics (#4596)
On an invalid ov.conf the server exited 1 with only the raw parse error. Bootstrap now reports the resolved config path and points the operator at 'openviking-server doctor' and examples/ov.conf.example. Also include the resolved path in the local 'server' section type error for consistency. FileNotFoundError output and the exit code (1) are unchanged.
2026-09-02 19:17:41 +08:00
Jiajie - He/him/his da2aaa0ed0 fix(memory): keep dynamic URI values in one path segment (#4587) 2026-09-02 15:55:05 +08:00
Qin Haojie 26a460060d perf(admin): batch account directory initialization (#4574) 2026-09-02 15:06:06 +08:00
chenyu 7b0482576c fix(language): ignore bare import paths during detection (#4543)
* fix(language): ignore bare import paths during detection

Exclude bare domain paths from language detection so repeated .com imports are not counted as Portuguese stopwords. Preserve the original paths in summarization prompts.

仅在语言检测时排除裸域名路径,避免重复的 .com import 被计为葡萄牙语常用词;摘要 Prompt 仍保留原始路径。

* test(language): cover import path noise in directory overviews

Cover code-only, Chinese, Portuguese, and explicit-override cases while ensuring import paths remain in overview prompts.

覆盖纯代码、中文、葡萄牙语和显式语言配置,并验证目录摘要 Prompt 保留原始 import 路径。
2026-09-02 11:31:52 +08:00
Leoy 6271dbbc74 fix(cli): make resource URI segments Windows-safe like memory uris (#4517 precedent) (#4549) 2026-09-02 11:21:39 +08:00
Qin Haojie c486088ee7 fix(semantic): skip contended parent freshness updates (#4559)
* fix(semantic): isolate parent freshness updates

* refactor(semantic): trust queue manager lifecycle

* refactor(semantic): make parent freshness best effort

* refactor(semantic): keep parent refresh in existing scope

* fix(semantic): wait briefly for parent freshness lock
2026-09-01 19:22:10 +08:00
Yu ZhangandClaude Sonnet 4.6 f2599b3a04 fix: preserve memory metadata in batch writes (#4386)
Route batched content writes through the same in-place writer as single-file writes so memory replace operations keep existing hidden fields.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-09-01 17:16:25 +08:00
Jiahui ZhouandTRAE CLI 4f9840eca7 feat(tags): support tagged write and filesystem filtering (#4457)
Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-09-01 16:22:35 +08:00
baojun-zhang 63c25306af feat(s3): add aliyun oss vendor support and config (#4564) 2026-09-01 15:13:18 +08:00
dingbenandTRAE CLI 3b83072f5c fix(admin): list accounts/users in creation order (#4554)
Restore creation (insertion) order for admin list_accounts/list_users,
reverting the always-on lexicographic sorting introduced in #4409. No
sorting parameters are exposed; listing simply preserves the order in
which accounts/users were registered.

Applied across the server manager/router, the Python/Go/TS SDKs, the ov
CLI, and the admin docs/tests.

Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-09-01 14:00:49 +08:00
LeoFreeandliuyuxin 65f60beb48 fix(mcp): preserve parent for local file uploads (#4521)
Co-authored-by: liuyuxin <uyuc@liuyuxindeMacBook-Air.local>
2026-09-01 12:20:21 +08:00
bot-of-qin-ctxandqin-ctx ed4bb192d8 fix(session): keep token rebuild off event loop (#4557)
* fix(session): keep token rebuild off event loop

* fix(session): offload all pending token rebuilds

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-09-01 12:14:38 +08:00
jasonzhang 78962c32ad fix(memory): redact inline image bytes from extraction (#4456) 2026-09-01 10:56:52 +08:00
zgy c73d7e3cc7 fix(session): use session skill schema for patch merge (#4395) 2026-08-31 21:18:18 +08:00
bot-of-qin-ctxandqin-ctx 02f5c46419 fix(vlm): reject unsupported stream configuration (#4541)
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-31 20:52:32 +08:00
Terminator666666andTerminator666666 ae50736ece fix(rerank): validate credentials for explicitly configured providers (#4443)
* fix(rerank): validate credentials for explicitly configured providers

RerankConfig checked required fields for openai and litellm only. An explicit
`provider: cohere` without api_key, or `provider: vikingdb` without ak/sk, was
accepted at load time. is_available() then returned False and
HierarchicalRetriever fell back to plain vector search, logging a single
info-level line saying rerank was not configured.

The new checks run against the effective provider, matching the existing
openai and litellm branches. Auto-detection is unaffected, since detecting
cohere already requires api_key and detecting vikingdb already requires ak and
sk. An empty RerankConfig() still resolves to no provider and stays valid.

Drops test_default_provider_is_vikingdb, which asserted a default that
auto-detection replaced and had been failing on main. Rewrites
test_unknown_provider_raises_value_error to actually cover an unknown provider
and adds coverage for the two providers that were missing validation.

* docs(configuration): state required credentials per rerank provider

The rerank section described credential inference but not the fields each
provider requires when provider is set explicitly.

---------

Co-authored-by: Terminator666666 <Terminator666666@users.noreply.github.com>
2026-08-31 20:16:37 +08:00
Onefly e73c5bfc97 fix(retrieval): keep VikingFS.find on quick mode (#4472) 2026-08-31 19:54:08 +08:00
Onefly e24b834e5a fix(semantic): do not materialize missing DAG roots (#4474) 2026-08-31 19:51:56 +08:00
Rami d5bd4fd7a8 fix(session): pass memory_policy through CreateSessionOptions (#4495)
AsyncHTTPClient.create_session() takes (session_id, options); memory_policy
is a key of CreateSessionOptions, not a keyword argument. Three production
call sites still pass it as a keyword and raise

    TypeError: AsyncHTTPClient.create_session() got an unexpected keyword
    argument 'memory_policy'

on every session that does not already exist:

- openviking/ingest/replay.py, ConversationReplayClient.ensure_session:
  `ingest backfill` fails on every new session. The orchestrator catches
  per-session exceptions, so a first backfill prints one error per session
  and finishes with 0 commits.
- bot/vikingbot/openviking_mount/ov_server.py, VikingClient.ensure_session.
- openviking/session/train/components/session_commit.py,
  SessionCommitPolicyTrainer._commit_one, which swallows the TypeError and
  returns a failed commit record with an empty task_id.

All three now pass options={"memory_policy": policy}, and options=None when
no policy is configured. benchmark/locomo/vikingbot/import_to_ov.py already
used that form.

The three test fakes accepted the obsolete keyword, so none of the paths had
regression coverage. They now mirror the real SDK signature: reverting any
one of the three fixes fails its tests.

Fixes #4493
2026-08-31 19:46:11 +08:00
bot-of-qin-ctxandqin-ctx cb7fe2ad8d refactor(memory): remove unused URI allowlist (#4536)
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-31 19:27:06 +08:00
ShaoZegangByte 0b3b407c9d fix(parser): handle angle-bracket image paths in Markdown (#4533) 2026-08-31 19:14:25 +08:00
ShaoZegangByte 4845ae39e3 Understanding directory (#4439)
* fix: directory understanding_api & fix temp empty file

* fix: understanding_api error处理

* fix: api err msg

* fix: html test

* fix: support directory and HTML imports via UnderstandingAPI

* fix: support directory and HTML imports via UnderstandingAPI

* fix: director max files and zip msg

* fix: director max files and zip msg
2026-08-31 17:28:59 +08:00
Qin Haojie 170e17c145 feat(acl): add account-level authorization switch (#4527) 2026-08-31 17:17:14 +08:00
Qin Haojie 4d00feac93 fix(acl): remove automatic schema migration (#4522) 2026-08-31 16:09:05 +08:00
DuTao 76ab53ac67 fix(memory): make generated URIs Windows-safe (#4517) 2026-08-31 15:50:20 +08:00
50228a2e70 feat: support vector record IDs across filesystem APIs (#4442)
* feat: support vector record IDs across filesystem APIs

- centralize deterministic vector record ID generation and migration handling
- allow stat and read to resolve file IDs with actionable missing-index diagnostics
- add selectable filesystem fields and script-friendly ls, tree, and glob output
- preserve complete IDs in simple output and honor tree --simple without fields
- restrict Python SDK ID passthrough to supported read-only endpoints
- preserve explicit empty glob extra_fields for metadata responses
- retain the dedicated Codex OAuth doctor diagnostic path
- document the public API behavior and add CLI, SDK, storage, and server regressions

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* fix(cli): preserve full record IDs in field tables

Render filesystem record IDs without abbreviation in both table and simple field modes so the values can be passed directly to stat and read. Add a regression test for normal table rendering.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: Maojia Sheng <shengmaojia@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-08-31 14:32:40 +08:00
DuTao 6248d4e41c fix(memory): Event memory writes reusing existing page IDs (#4388)
* 修复 event paget id错误,导致未新建的问题

* fix(memory): enforce response-scoped page id uniqueness

* fix(memory): reject relations from mismatched page ids
2026-08-31 11:53:20 +08:00
txfangandchrisfang 3123e8d8ff feat(ragfs): unify CacheRuntime and Redis-backed CacheFS/QueueFS (#4353)
* feat(ragfs): unify cache runtime providers

* refactor(ragfs): remove bundled external cache providers

* refactor(ragfs): unify cache runtime and redis backend

* chore: sync contributing guide with main

* fix(cache): harden Redis runtime compatibility

* refactor(cache): gate memory mock behind test feature

* fix(cache): harden Redis runtime integration

---------

Co-authored-by: chrisfang <chrisfang@noreply.gitcode.com>
2026-08-31 11:10:08 +08:00
Jiahui ZhouandTRAE CLI 550ef79675 feat(temp-upload): hour-bucketed shared uploads with best-effort cleanup (#4479)
Group shared uploads into hourly buckets so the upload root never becomes one
huge flat directory (which made cleanup enumerate hundreds of thousands of
entries and take minutes), and rework cleanup to be bounded and best-effort.

Storage layout:
- New uploads: viking://upload/<YYYYMMDDHH>/<uuid>/{content,meta}, UTC hour.
  upload_id is <YYYYMMDDHH>-<uuid>; the 10-digit hour prefix is the bucket and a
  format marker, and the in-bucket directory is just the uuid (no repeated
  prefix). temp_file_id stays shared_<upload_id> (external contract unchanged).
- Legacy flat <13-digit-ms>-<uuid> uploads stay readable; _read_shared_meta
  picks exactly one path by id format, no fallback probe.

Cleanup (best-effort, off the request path, oldest-first via name-ascending ls):
- module-level due_at (epoch) + pending throttle so requests don't pile up
  duplicate jobs;
- YYYYMMDDHH bucket expires at bucket_start + 3600 + ttl and is removed whole
  (rm -r); scan stops at the first live bucket;
- legacy flat uploads expire by created_at + ttl and are removed individually;
- other/malformed dirs are removed only when temp_upload.cleanup_invalid_dirs is
  enabled, and legacy flat uploads are never treated as invalid;
- every deletion logs kind/uri/elapsed_ms.

Shared upload writes, failure rollback, and cleanup deletes use
auto_pathlock=False; VikingFS.rm/write_file/write_file_bytes gain an
auto_pathlock parameter (default True, no behavior change for other callers).

Add temp_upload.cleanup_invalid_dirs config flag (default False).

Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-08-31 11:08:01 +08:00