The api_key_watch_interval mechanism reloads account info across replicas
by polling compute_store_signature() (a (path, size, modTime) signature
over accounts.json + users.json). On S3FS this never detected writer-side
changes: S3FS stat() serves from a sliding-TTL StatCache (default 60s,
re-armed on every get()), so the watcher's 30s poll kept the entry alive
forever and the signature never moved -> reload() never fired. localfs
stats live, so it worked there.
Thread a bypass_cache flag from the watcher's stat call down to S3FS so
signature stats read fresh backend metadata:
- FsContext(View): add bypass_cache field + builder/getter
- ragfs-python build_fs_context: parse ctx["bypass_cache"]
- S3FS stat(): skip stat_cache.get() when bypass_cache is set
- AsyncAGFSClient.stat(bypass_cache=...): inject ctx flag
- legacy _stat_signature: stat with bypass_cache=True
Co-authored-by: TRAE CLI <traecli@bytedance.com>
The public MountableFS constructors do not initialize a pathlock manager, but
multi-write mount and raw copy used expect() on the missing manager, so calling
either fast path on such an instance panicked.
Return Error::Config with a clear message instead; tests that need the success
path already build the manager via with_test_pathlock_manager().
* chore: remove dead git tuning knobs and duplicate release/frontend files
- GitTuningConfig: drop upload_concurrency, restore_concurrency,
ref_cas_max_retry, and ref_cas_backoff_ms, which were parsed but never read
anywhere (verified no source readers). Keep commit_index_enabled and
blob_exists_precheck_enabled, the two knobs that actually take effect.
- Design doc: mark the removed knobs as roadmap items to be added back when
the behavior lands, and correct stale 'not implemented' claims
(validate_account_id is enforced at the Git service entry; blob reads are
limited via show_with_limit).
- Remove .github/workflows/release-vikingbot-first.yml: a historical one-off
PyPI release workflow whose bot/ package lacks build metadata.
- web-studio: delete pnpm-lock.yaml and the pnpm-only package.json block;
the Makefile and CI already use npm + package-lock.json as the single
install chain. Net -9,695 lines.
* docs: align cleanup notes with current behavior
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
Replace the fixed 50 ms retry loop with bounded exponential backoff and jitter while preserving the original first-retry latency floor. Clamp each delay to the remaining timeout so retry sleeps do not add avoidable timeout overshoot.
Add deterministic interval and polling-reduction checks plus a four-waiter contention regression. The 10-second worst-case schedule drops retry probes from 200 to 29 while all waiters still acquire after the holder releases.
Tests: cargo test -p ragfs --lib lock:: -- --nocapture
* fix(pathlock): treat localfs read ENOENT as NotFound
Map ENOENT from localfs read back to NotFound when a file disappears
between metadata() and fs::read().
This avoids misclassifying lock-file delete races as plugin I/O errors
in pathlock token reads, so missing .path.ovlock is handled as an
expected absence instead of a fatal lock I/O failure.
* fix(pathlock): treat localfs read ENOENT as NotFound
Map ENOENT from localfs read back to NotFound when a file disappears
between metadata() and fs::read().
- keep handed-off leases in the registry with pending_handoff so auto-refresh continues
- include lease_ref in PathLockHandoffRef for same-process adopt fast path
- rotate both lease_ref and ownership_ref on adopt to invalidate stale producer capabilities
- reject stale capability use in release, release_selected and refresh
- validate owner_id, lock_paths and covered_paths before local adopt
- reject replayed fallback adopt when the same owner/path token is already held
- add tests for pending handoff refresh, retryable adopt race, forged coverage and replay rejection
* fix(ragfs): preserve cache visibility on partial S3 deletes
Surface exact and per-object S3 deletion failures, while always invalidating the affected directory and stat cache scope after a recursive delete attempt.
Source-PR: #3407
Original-Commit: 8d6addf28e
* fix(session): preserve legacy policy and peer identity compatibility
Parse string false and other legacy boolean-like memory policy values without silently enabling extraction or breaking persisted configs. Encode mixed-script peers losslessly, while retaining their former lossy IDs as read-only retrieval and extraction aliases.
Source-PR: #3422
Original-Commit: 0dfd5a9ed9
* fix(memory): drain timer flush tasks during shutdown
Retain the shielded timer flush task and await it when close cancels the timer loop, so batch failures are observed and submitters are resolved without unhandled task exceptions.
Source-PR: #3438
Original-Commit: ca1d74e164
* fix(storage): preserve peer isolation and cache correctness
* fix(ingest): reserve encoded peer namespace
* ci: skip embedding-dependent resource test without secrets
* fix(ragfs): invalidate caches after partial remove
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
- add storage.agfs.pathlock.lock_timeout_secs
- use pathlock default timeout instead of hardcoded zero in wrapper
- map legacy storage.transaction.lock_timeout when new config is unset
- remote redolog by using persistent `session_commit` queue.
* feat: implement server-resolved OpenViking Assets manifests
Add the openviking-assets/1 declaration flow with server-owned configuration parsing and native Rust CLI execution.
- Resolve one flat Manifest against one Catalog through an authenticated server endpoint with strict schema and Git semantic validation.
- Reject recursive includes and unsafe clone URLs; return a resolved plan without submitting resources or running server-side batches.
- Keep local credential aliases, manifest state, dry-run, failure isolation, and per-asset create/sync orchestration in the CLI.
- Generate normalized stable asset identities on the server and remove the CLI direct SHA-1 dependency.
- Update flat examples and add server resolver/API plus Rust CLI coverage.
* feat: implement server-resolved OpenViking Assets manifests
* feat: implement server-resolved OpenViking Assets manifests
* fix(pathlock): tolerate missing lock token after recursive delete
* feat: implement server-resolved OpenViking Assets manifests
* feat: implement server-resolved OpenViking Assets manifests
* fix(ragfs): enforce mount containment and write-flag semantics in LocalFS
Sweep findings: E-01, E-04. Reject lexical traversal and honor LocalFS write contracts.
* fix(review): reject absolute localfs remainders
Addresses blocking review finding on #3402.
* fix(ragfs): close LocalFS glob mount-escape gap
LocalFS::glob_directory only called validate_virtual_path(path), which
rejects `..` but accepts an absolute remainder. A `//`-double-slash mount
path (e.g. `/local//etc`) survives normalize_path, and find_mount hands the
remainder `//etc` to glob_directory; validate_virtual_path passes it
(components are [RootDir, Normal("etc")], no ParentDir), and glob_via_walk's
resolve_virtual_path strips one slash to `/etc` and joins it over the base —
an absolute join that overrides the base and lists the host directory.
Every other op (read/write/stat/rename/remove/grep) routes through
resolve_path, which adds the is_absolute check. glob now calls the same
resolve_path guard (discarding the returned PathBuf) so it shares the
containment contract. Directory-listing info leak only; content reads were
already covered.
Adds test_localfs_glob_mount_rejects_absolute_remainder mirroring the read
regression test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RbWx1T81KkNV4nucxWsPXZ
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(storage): optimize glob func
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* fix(localfs): offload blocking fs operations to spawn_blocking
* feat(glob): cap glob api default node_limit at 256
* feat(sdk): add node_limit options for glob in python and go SDKs
Fixes unbounded queue.db growth for newly created QueueFS SQLite
databases by enabling SQLite auto_vacuum=FULL at DB initialization.
- Detect brand-new databases (no tables in sqlite_master)
- Set PRAGMA auto_vacuum=FULL and run VACUUM before schema creation
- Preserve existing deployed databases unchanged (no risky rewrites)
- Add regression tests for both new-db and legacy-db behavior
Fixes#2707
* add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
* Handle poisoned known_keys mutexes in cache providers
* Add configurable cache support to ragfs python
* Add Redis cache provider support
* Prune native cache providers from default build
* Add RAGFS cache guides in Chinese and English
* delete .cargo/config.toml
* Fix stale cache invalidation across shared wrappers
* Document Yuanrong native concurrency limits
* add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
* Handle poisoned known_keys mutexes in cache providers
* Add configurable cache support to ragfs python
* Add Redis cache provider support
* Prune native cache providers from default build
* Add RAGFS cache guides in Chinese and English
* delete .cargo/config.toml
* Fix stale cache invalidation across shared wrappers
* Document Yuanrong native concurrency limits
* Limit directory cache entries and add regression test
* docs: add TOS s3fs backend test design
* Add runtime cache config override
* add build guides for mooncake and yuanrong
* fix _FakeConfig test with cache
* fix(ragfs): cache encrypted data below the encryption layer
- pass runtime cache config to the Rust binding
- cache ciphertext instead of decrypted content
- disable cache for encrypted multi-write mounts
- update tests and provider build documentation
* fix(ragfs): invalidate cache after same-mount raw copy
Ensure copy_within_mount invalidates the cached destination file and
parent directory after the raw backend fast path writes data directly.
This prevents stale reads and stale directory metadata when cache is
enabled and Python cp() uses the same-mount copy optimization.
Add a regression test covering overwrite copy followed by read/list.
* fix(ragfs): preserve multi-write discovery through cache layer
Allow MountableFS::as_multiwrite() to unwrap CachedFileSystem when
discovering the underlying MultiWriteWrappedFS. This preserves
multi-write admin paths and same-mount copy behavior for cache-enabled,
unencrypted multi-write mounts.
Add a regression test covering sync status, sync retry, same-mount copy,
and unmount behavior for cached unencrypted multi-write mounts.
* feat(ragfs): add cache-aware tree traversal mode
- add configurable tree traversal mode to cache policy
- keep default tree behavior delegated to backend
- allow cached traversal to reuse read_dir directory cache
- bypass cached traversal for multi-write backends
- add regression coverage for tree cache behavior and fallbacks
* docs: design cache-aware grep traversal
* feat(ragfs): add cache-aware grep traversal
Introduce a shared cache traversal mode for recursive APIs and use it to
optionally run grep through CachedFileSystem.
- add CacheTraversalMode with backend and cached_traversal modes
- keep CacheTreeMode as a compatibility alias
- route tree and grep through cached traversal only when explicitly enabled
- reuse cached read_dir entries and full-file reads during grep traversal
- keep multi-write traversal on the backend path
- expose storage.agfs.cache.traversal_mode in Python config
- raise max cached directory entries threshold to 4096
- add regression tests for grep cache traversal and traversal config
* Optimize cached grep generation validation
* Parallelize cached grep file scanning
---------
Co-authored-by: fang <fang@fangMacBook-Air.local>