* fix(resource): reject a source with no content at ingestion
Adding a zero-byte file currently succeeds and produces an empty
resource. With the Understanding API disabled -- the default, since
ParserApiConfig.enable is False -- nothing on the internal parse path
looks at the size, so the file is staged, parsed, and indexed as an
empty entry. Directory imports already refuse the same input:
directory_scan.py skips any zero-byte member as an "empty file". A
single-file import should not disagree with that.
Check it in UnifiedResourceProcessor.prepare, the one point every
ingestion path passes through: prepare_durable_source freezes a source
there before the request touches the tree, and process() calls it for
anything not frozen earlier. So local files, uploads, remote downloads,
git and Feishu sources are all covered by one check, at the moment the
bytes are first in hand.
InvalidArgumentError maps to 400 through ERROR_CODE_TO_HTTP_STATUS.
Where ingestion runs asynchronously -- a plain remote URL with the
Understanding API off is queued, and the response has already been
sent -- the same error fails the task instead of the request.
Deliberately narrow:
- Only zero bytes. A one-byte file is still accepted; this is not a
minimum-size policy.
- Only regular files. A directory has no meaningful size and is skipped,
so directory and repository imports containing empty files are
unaffected.
- A stat failure is left to the normal ingestion path to report.
- content/write is untouched. Creating an empty file there is an explicit
user action, not an ingestion accident.
The error names the caller's own file rather than the temp working copy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(resource): clean rejected temporary sources
* fix(resource): preserve queued error codes
* test(resource): focus empty-source regression coverage
* ci: drop dedicated empty-resource regression step
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(uri)!: reject uid-less current-user shorthand in favor of viking://~
viking://user/<segment> (memories/resources/skills/peers/privacy/sessions
without a user id) was ambiguous with a user literally named after the
segment, and a user actually named e.g. "memories" was unreachable for
USER/ADMIN callers. Now that the viking://~ home alias (#4167) covers the
same need unambiguously, the shorthand fails closed at the request
boundary instead of expanding:
- resolve_current_user_uri raises NamespaceShapeError with a corrective
hint naming both viking://~/<rest> and the explicit-uid form. Silently
parsing the reserved segment as a peer user id would misdirect reads
and writes, so rejection is the only safe removal.
- Bare viking://user falls through to the canonical parser and keeps
container semantics (a user key listing it sees only its own space).
- The self-id escape stays: a caller whose user_id equals a reserved
name keeps viking://user/<own-id> as their canonical root. ROOT-role
literal parsing and the legacy viking://session alias are unchanged.
- AddTargetsConfig normalizes stored legacy config spellings
(viking://user/resources|skills) to the viking://~ form at validation
so existing ov.conf/user_config deployments keep working; the accepted
per-user spelling is now viking://~/resources and viking://~/skills.
- usage_reporter keeps canonicalizing the historical shorthand found in
old transcripts and additionally recognizes viking://~/memories/.
BREAKING CHANGE: requests using the uid-less viking://user/<segment>
spelling now fail with 400; use viking://~/<segment> or an explicit
viking://user/{user_id}/<segment> URI.
* refactor(clients): migrate first-party emitters to the viking://~ home alias
Every in-repo client that emitted the removed uid-less current-user
shorthand now sends viking://~/... instead: vikingbot fallbacks and
default sentinels, the LangChain store/tools defaults, the shared
recall-core.mjs (all synced plugin copies), the codex/claude-code/
openclaw/openwebui/dsh/zcode/pi plugin emitters, quick-app examples,
Go SDK example, tau2 benchmark targets, and the eval golden dataset.
Compat kept where legacy strings live in stored user configs: bot and
ov_dream sentinels accept both spellings while emitting only ~, and
recall-core still rewrites legacy viking://user/<reserved> config values
client-side. langchain_openviking._uri now classifies viking://~ with
the explicit-user shape so canonicalized server responses keep matching
a ~ root. Plugin READMEs note the server requirement for the alias.
* docs: replace current-user shorthand guidance with the viking://~ home alias
Rewrite every EN/ZH doc and model-facing prompt that advertised the
uid-less viking://user/<segment> spelling: URI concept catalogue,
context-types/storage/extraction/retrieval/session/privacy concepts,
configuration guide (with the legacy add_targets auto-normalization
note), resources/skills/sessions/retrieval/admin API references, FAQ,
capability reference, and the openviking-memory / ov-experience-memory /
openclaw / ov-resources skills. The stale MCP viking://user/<path>
dialect passage in the MCP guide is replaced by ~ guidance, and bare
viking://user is documented as the container of user spaces.
* test(api): migrate live API session-used tests off the removed shorthand
tests/api_test/sessions sent uid-less viking://user/skills/... URIs to
record_used, which the request boundary now rejects with 400 (caught by
the API & CLI Integration Tests CI job; these tests need a live server
and are not part of the local suites). The api_test client authenticates
as an admin-role user key, so the viking://~ home alias expands for it.
tests/api_test/common/test_edge_cases.py is left as is: it asserts a 400
for a non-resource add target, which still holds.
* feat(storage): optimize glob func
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* fix(localfs): offload blocking fs operations to spawn_blocking
* feat(glob): cap glob api default node_limit at 256
* feat(sdk): add node_limit options for glob in python and go SDKs
* fix: stabilize studio identity and streaming chat
* fix: hide unsupported studio terminal commands
* fix: remove unsupported terminal command copy
* fix: run selected terminal suggestion on enter
* fix: group supported terminal commands
* fix: add terminal quick start and history
* fix: scope session visibility by user
* fix: harden bot user scoping
* fix: forward request scoped bot identity
* fix: add terminal quick start translations
* fix: add terminal command group translations
* fix: simplify studio identity scoping
* fix: support api key copy on dev urls
* fix: stop passing agent id to ov http client
* fix: search follow-up memory questions
1. Skip find/add_skill tests on 401/500 when embedding unavailable
- PR CI (fork repos) cannot access VLM/Embedding API keys, causing
embedding service to return 401/500 with dummy keys
- test_response_types.py: skip test_find_response_types on 401
- test_skill_api.py: skip add_skill tests on 500, find tests on 401
- build_test_helpers.py: skip assert_resource_findable on 401
2. Run CLI tests serially to avoid CONFLICT and timeout
- CLI tests share session-scoped fixtures (test_dir_uri, test_pack_uri)
- Concurrent workers cause CONFLICT Resource is busy and timeout errors
- Remove -n 4 from CLI compatibility and integration test steps
3. Reduce parallelism in effect tests to avoid server overload
- Lightweight tests: -n 4 -> -n 2
- Heavy tests: -n 2 -> serial (resource-intensive operations)
- Add CLI integration tests (10 test files under tests/cli/)
- Extend api_test.yml with CLI install + test steps
- Run filesystem + scenarios/resources_retrieval serially to avoid 409 conflicts
- Other tests parallel with -n 4
- Add release prereleased trigger to api_test.yml and api_test_effect.yml
- Deduplicate oc2ov_test P0 cases (20→12, ~30-55min saved):
- Delete test_memory_write.py (covered by V2 suite)
- Remove events/tools from V2 suite (structurally identical to entities/skills)
- Remove test_memory_read_verify (covered by V2 suite)
- Remove test_cross_session_recall (overlaps with recall_explicit_search)
- Add ensure_resources_dir fixture to prevent NOT_FOUND on fresh environments
- Add retry logic for 429/500/403 rate-limit in api_client.py
- Add retry for commit when task_id is None in test_memory_v2_full_suite.py
- Add exponential backoff retry for GitHub platform test 5xx errors
* chore(format): align python and c++ file formatting
* chore: update urllib3 to 2.7.0 and clean test imports
1. bump urllib3 dependency from 2.6.3 to 2.7.0
2. remove unused pytest import and RoleScope import from test file
* style: format list comprehensions and lambda function for readability
Adjust the line breaks in the list comprehension in the VikingSearchTool class to follow standard Python formatting conventions, and rewrap the lambda assignment in the test case to improve code readability without changing functionality.
* style: fix line wrapping and remove extra blank line
- remove stray blank line in ov_server.py
- wrap long logger.info line in memory.py for better readability
* style: fix targeted ruff lint violations
* chore: clean up unused imports and reorder code
This commit removes unused imports, reorders import statements for better consistency,
and simplifies some test file imports. Changes include:
- Remove redundant blank lines and unused imports across multiple test files and core modules
- Reorder imports in openviking hooks module to follow standard layout
- Fix import ordering in memory isolation handler
- Simplify php parser type imports
- Move volcengine mock import to correct position in test file
* refactor(uri utils): remove extra blank lines in uri.py
clean up redundant whitespace to improve code readability
* feat(ovpack): add v2 manifest and conflict policy
Add a portable OVPack manifest for scalar metadata and make imports validate scope, derived files, and conflicts before writing.
* fix(ovpack): remove import vectorize option
Make OVPack imports always rebuild vectors in the target environment, keep legacy packages compatible, and reject unsupported manifest versions before writing.
* fix(ovpack): remove force import alias
Use on_conflict as the single OVPack import conflict policy and reject removed force inputs.
* fix(ovpack): regenerate runtime vector metadata
Keep type portable but stop exporting or applying created_at, updated_at, and active_count from OVPack manifests.
* fix(ovpack): validate manifest contents
* fix(ovpack): require manifests for imports
* fix(ovpack): close manifest validation gaps
* fix(ovpack): defer parent creation until validation passes
* fix(ovpack): remove export size guard
* fix(ovpack): support session and scope-root restores
* docs(ovpack): document full backup migration
* feat(ovpack): add backup restore workflow
* fix(ovpack): validate import scope compatibility
- test_build_media_resources_slow: assert 500 for SVG parse failure
- test_build_platform_wikipedia: assert 500 for Wikipedia URL fetch failure
- test_build_error_handling_slow: assert 500 for corrupted ZIP (align with non-slow version)
- test_memory_v2_full_suite: add find_session_by_id fallback for CI deterministic session IDs
After #1744 (fix(api): return processing errors as error envelopes),
the server returns proper HTTP error status codes instead of always 200:
- Corrupted ZIP: returns HTTP 500 with PROCESSING_ERROR envelope
(was HTTP 200 with inner status='error')
- Invalid URI (local paths like /tmp/...): returns HTTP 400 with
INVALID_ARGUMENT envelope (was HTTP 200 with inner status='error')
Changes:
- test_build_error_handling.py: assert status_code == 500 and
response status == 'error' for corrupted ZIP
- test_fs_mv.py: use viking://resources/ URIs instead of /tmp/ paths,
add cleanup in finally block
- test_fs_rm.py: use viking://resources/ URIs instead of /tmp/ paths
After #1744 (fix(api): return processing errors as error envelopes),
the server returns HTTP 500 with a structured error envelope for
processing errors like corrupted ZIP files, instead of HTTP 200.
Update test_error_corrupted_zip to assert status_code == 500 and
response status == "error".
Wikipedia blocks requests from cloud datacenter IPs (Azure/GCP/AWS),
causing 403 Forbidden in GitHub Actions. Add graceful handling:
- Check outer/inner error for 403/forbidden/blocked keywords
- Print skip message and return instead of hard failure
- Still validates full flow when Wikipedia is accessible
* fix(security): clean up code scanning and runtime findings
Harden path and logging boundaries, remove noisy cleanup issues,
and keep observability failures from breaking runtime flows.
* fix(security): close werewolf and feishu validation gaps
Block the remaining path traversal bypass in the werewolf demo,
and validate Feishu hosts on the main parse() entry point.
* feat(session): add account namespace policy and shared sessions
Unify namespace resolution across filesystem, indexing, and session storage.
Add account-shared session paths, role_id auth semantics, and an HTTP demo
script for the four namespace-policy combinations.
* space
* fix(pack): skip derived semantic files in ovpack transfer
Keep ovpack imports resilient to stale sidecars and rebuild semantics through the normal queue instead of restoring derived files verbatim.
* Revert "fix(pack): skip derived semantic files in ovpack transfer"
This reverts commit f4e4db8401.
* fix(namespace): default legacy accounts to agent-shared policy
Clarify that memory.agent_scope_mode is deprecated and document the supported agent memory migration paths.
- Add pre-start cleanup in upgrade_openviking.sh: remove stale .lock
and .jsonl.lock files, kill all residual openclaw processes before restart
- Add session lock awareness in openclaw_cli_client.py: wait for lock
release before sending requests and after subprocess returns
- Optimize wait_for_sync in base_cli_test.py: enforce minimum 5s wait,
check lock release before proceeding, raise poll interval to 3s minimum
- Add Chinese descriptions for Memory V2 test cases in conftest.py
- Fix test_session_commit assertion: remove pre_archive_abstracts check
since get_session_context no longer returns this field
* fix: make api_test more robust for CI environments
- Add @pytest.hookimpl(optionalhook=True) for pytest-html hooks to fix compatibility issues
- test_fs_read: skip test when AGFS service is not available
- test_get_overview: skip test when overview file does not exist
These changes ensure tests pass gracefully on CI servers where AGFS service
may not be available or files may not exist.
* ci: reduce max-parallel to 1 for better resource availability
Reduce max-parallel from 2 to 1 to avoid waiting for multiple runners
when GitHub-hosted runners are limited.
add support for trailing slash rules in resource URIs to control file/directory placement
update CLI, API, and documentation to reflect new URI handling semantics
add comprehensive tests for all URI semantics cases
* feat: add scenario-based API tests
- Add scenario test framework with proper categorization
- Add tests for resources_retrieval, sessions, and stability_error scenarios
- Add get_task() and wait_for_task() methods to API client for async operations
- Add get_session_context() method for session context retrieval
- Update API test workflow name from '03' to '06'
- All 14 scenario tests pass with proper business logic validation
* fix: skip scenarios tests when no VLM/Embedding secrets
Scenarios tests require VLM for session archival summaries and
Embedding for semantic search. Skip them in basic test mode.
* fix(bot): respect OPENVIKING_CONFIG_FILE even when file doesn't exist
Previously, bot's config path resolution would fallback to
~/.openviking/ov.conf when OPENVIKING_CONFIG_FILE was set but the
file didn't exist. This was inconsistent with server's behavior,
which treats a missing env-specified config as an error.
In container deployments with OPENVIKING_CONFIG_FILE=/app/ov.conf,
this caused bot to potentially write auto-generated config to
/root/.openviking/ov.conf instead of the intended /app/ov.conf,
leading to config file path mismatches.
Now bot respects the environment variable unconditionally, matching
server's behavior and ensuring both components use the same config
path.
Fixes#1242🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(embedder): support dimension truncation for OpenAI-compatible models
Previously, when users configured a dimension (e.g., 1024) for OpenAI-compatible
embedding models that don't support the 'dimensions' API parameter, the system
would fail with "dimensions is currently not supported" error.
Additionally, when dimension was not configured, the config layer would return
a hardcoded fallback of 2048, but the actual model might return a different
dimension (e.g., 1024), causing dimension validation failures.
This fix implements vector truncation in OpenAIDenseEmbedder:
- Removes the 'dimensions' parameter from API calls (not supported by all models)
- If user configures dimension=1024 and model returns 2048, truncates to 1024
- If no dimension is configured, uses model's native dimension without truncation
- Applies truncation to both single and batch embedding operations
This allows users to:
1. Use OpenAI-compatible models with custom dimensions via truncation
2. Control vector dimensions for storage optimization
3. Avoid dimension mismatch errors between config and actual embeddings
Fixes#1238🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: https://github.com/volcengine/OpenViking/issues/1232
---------
Co-authored-by: openviking <openviking@example.com>
Co-authored-by: Claude <noreply@anthropic.com>
* feat(sessions): support specifying session_id when creating session
* feat(session): validate session_id uniqueness on create
Add AlreadyExistsError check in session_service.create() when a specific
session_id is provided, ensuring idempotent behavior and preventing
accidental overwrites of existing sessions.
* fix(http): replace temp paths with upload ids
Stop exposing server filesystem paths through temp uploads and require
HTTP callers to use temp_file_id across server, clients, tests, and docs.
* fix(cli): upload local ovpacks in http mode
Make the Rust HTTP client import local ovpack files through temp uploads
and cover the flow with an end-to-end SDK regression test.
* fix(api): align integration client with temp upload contract
* fix(client): fail fast for invalid ovpack imports