`list` was the only listing tool whose `uri` was required; `tree` and `glob`
already default to viking://, so a call without a uri failed validation
instead of listing the root.
When `edit` finds no match but the old_string would match after normalizing
CRLF/LF, the error now says so instead of only asking for a re-read.
Co-authored-by: wzc1753 <81469626+wzc1753@users.noreply.github.com>
* feat(acl): default shared resources to inherited manager access
Use immutable user:* manage at the shared root and inherit permissions without
extra creator grants. Tolerate temporarily different ACL index snapshots.
* feat(acl): unify permissions under attrs and support creation attributes
* refactor(acl): accept top-level ACL on resource creation
* fix(acl): bind import permissions to authorized targets
Reject internal ingestion options from public resource arguments and defer
ACL authorization until the final import target has been resolved.
* refactor(acl): isolate import permissions from parser arguments
Build ingestion options from explicit public inputs and keep parser arguments out
of post-processing so they cannot supply internal ACL updates. Remove redundant
mkdir ACL handling and make permission snapshot selection easier to follow.
* fix(acl): release import locks after permission failures
Release locally acquired leases when source commit fails or is cancelled,
and ensure artifact cleanup cannot skip post-processing lock release.
* fix(docs): use CLI labels in ACL API references
* refactor(acl): restore dedicated permission interfaces
Restore standalone ACL HTTP, CLI and SDK operations while keeping attrs
focused on its existing attributes. Retain creation-time ACL support and
align the documentation with the final permission model and interfaces.
* fix(acl): authorize connector imports before creating watches
Reject ACL changes before reserving a Watch so denied requests cannot leave it stuck executing. Cover denial and submission cancellation in the existing Watch cleanup test.
* fix(web-studio): align ACL guidance and restriction status
Describe inherited permissions without creator privileges in both locales. Show access restriction only for restricted mode and verify the existing toggle flow against inherited management grants.
Add clear semantics across resource ingestion, content writes, reindexing, RNFV scalar planning, SDKs, CLI, and Web Studio. Treat empty replace tags as a no-op while preserving clear through durable queues and only updating vector scalars when the normalized value changes.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat(observer): support JSON status format
Add ?format=json to HTTP observer APIs and keep table as the default response shape. Expose the observer format option through the Python, Go, and TypeScript SDKs while preserving backward compatibility for existing calls.
Also add targeted server and SDK regression coverage for format propagation and structured observer payloads.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* fix(observer): return structured json fallback status
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* fix(observer): align json status summaries
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Account listing only supported case-sensitive `name` wildcard matching,
while user listing already exposed a case-insensitive `query` substring
search. Mirror that behavior on `GET /api/v1/admin/accounts` so callers
can find accounts by a case-insensitive id fragment.
Also forward the existing `query` parameter through the Python SDK's
`admin_list_accounts`/`admin_list_users` (sync and async), which
previously dropped it so the server-side search was unreachable.
- server: get_accounts accepts query_filter (casefold substring), route
exposes ?query=, NewAPIKeyManager forwards it to the legacy manager
- sdk: thread query through both admin list methods
- docs: document the account query parameter (en/zh)
- test: extend test_list_accounts contract with a case-insensitive match
Co-authored-by: xiaojian.xj <xiaojian.xj@bytedance.com>
* fix(retrieve): return one entry per Skill package from context assembly
A Skill package carries one vector per file and one per directory level,
so a single package answers a context request with its .abstract.md, its
.overview.md, its SKILL.md and every script or reference inside it. The
skills bucket counted those records rather than packages: with the coding
preset's quota of 2, one package could take both slots and still render a
single entry, because the cross-bucket merge that collapses a directory's
two sidecars runs after the quota cut. The entry that survived pointed at
whichever file matched, described by that file's own summary, and a hit on
a different file next turn made the cross-turn ledger miss it.
Skill hits now become their package as they are built, before exclusion is
checked: the URI is the package's SKILL.md, the same path /skills/find
reports as skill_md_uri, and the entry is a file rather than a directory,
so it follows the category's abstract default instead of the directory
rule that reads .overview.md. The bucket merges those candidates before
its quota applies and asks retrieval for four times the quota, since most
of what comes back collapses. The text is the package's own abstract, read
once per entry; a package whose abstract is not generated yet degrades to
a bare URI rather than borrowing the matched file's summary.
Quota-free flat retrieval merges skills the same way. The other categories
keep their own collapse, which happens later and on different terms.
Introduced by #5045, which moved Skills to whole-package indexing. The MCP
find tool already collapses to packages through SkillPackageRetriever; the
context face, the path every harness plugin actually takes, did not.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
* fix(retrieve): let the skills bucket use package retrieval
The skills bucket asked the generic retriever for four times its quota and
merged the rows into packages afterwards, because a package answers with
one record per file and per directory level. Four times is a guess: a
package matching in more than eight records still leaves the bucket short.
find_skills already does this exactly, paging until it holds the requested
number of distinct packages, and covers both skill roots in one call.
Two of its defaults stopped it being usable from here:
- It built SkillPackageRetriever without a rerank_config, so its threshold
fell back to 0 while viking_fs.find, which passes one, uses the
configured 0.1. A caller that leaves score_threshold unset was getting a
different filter from each path.
- It took no filter, so a context request carrying one would have had it
silently dropped for this bucket. retrieve_skills already accepted the
scope_dsl that backs it.
Package retrieval embeds the query, so the bucket only reaches for it when
there is query text. A filter-only request stays on the generic path, which
recognises it and scans scalars instead; routing it to find_skills would
embed an empty string, which the OpenAI-family embedders reject outright.
An image-only request has nothing either path can use here: an image query
forces the resource context type, which the skill filter then excludes, so
the bucket is skipped rather than searched.
A request that carries both a query and an image now gets skills back for
the text, where before it always got none.
Package normalization, the pre-quota merge and the package abstract stay
where they were. find_skills returns one hit per package but keeps the
matched file's URI and summary, and planned queries still overlap.
The rerank_config also reaches REST /api/v1/skills/find, whose
score_threshold defaults to None: its floor moves from 0 to the configured
0.1 and matches the generic path. Same for `ov skill find` and the SDK,
which both leave it unset. MCP find calls service.search.find, not this,
and is unaffected either way.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
SkillOperationUpdater creates a new skill with SkillProcessor.process_skill
but updates an existing one with a generic ContentWriteCoordinator.write of
SKILL.md. Under the user's own root that write is always refused, because
skills/ is one of the managed subtrees content_write.py rejects, so an
update only logs an error and the merged skill is lost. Under
viking://agent/skills the write lands, but the skill's .overview.md is then
regenerated by the semantic worker as a VLM summary instead of holding the
SKILL.md body the installer writes — and a skill's L1 vectors come from that
body, so the update degrades its own retrieval.
The tests replaced ContentWriteCoordinator with a fake that wrote the file,
so neither showed.
Both branches now install under the root the operation names, which also
fixes creation: it passed no root at all, so a session-derived skill
addressed to viking://agent/skills was created under the user's root. That
makes the two branches one call, so they are folded together.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* refactor(skills): install skills through one shared helper
POST /api/v1/skills kept its whole install loop (source resolution, per-skill
install, source metadata, list_only) inline in the route. Move it into
openviking/server/skill_ingest.py:install_skills so the MCP add_skill tool
and signed skill uploads can reuse the exact same code path. The REST
route's behavior is unchanged.
Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28
* feat(mcp): add an add_skill tool
MCP clients had no way to create a skill: write refuses the skills/
subtree (_USER_MANAGED_SUBTREES) and add_resource validates its target as
a resource. Agents that should keep skills in OpenViking could read them
but never add one.
add_skill takes either the full SKILL.md text (data) or a path. A Git or
GitHub tree URL installs through the same source resolution as REST, with
skills=[...] to pick from a multi-skill repository and list_only to
preview it. A local SKILL.md, directory, or zip gets the add_resource
treatment: the tool mints a one-time upload token, now tagged kind="skill"
with the target root, selection and list_only, and the signed temp_upload
installs the file as skills instead of ingesting it as a resource.
target_uri="viking://agent/skills" shares the skill with the account.
All three paths (REST, MCP inline/Git, signed upload) go through
skill_ingest.install_skills. The tool count in the server log, app
comment, docs, and the Codex plugin's REAL_MCP_TOOLS moves to 16.
Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28
* fix(mcp): search shared skills in find(context_type="skill")
Without a target_uri, find resolved the generic default targets, which
stop at the caller's user root, so a skill search never reached the
account-shared viking://agent/skills. REST /skills/find and the context
search already cover both roots. When context_type resolves to skill only
and no target_uri is given, the MCP tool now targets
default_target_directories(ctx, context_type=SKILL): the user's own skills
plus viking://agent/skills. REST find semantics are unchanged.
Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28
* fix(mcp): print directory abstracts in tree(include_abstract=true)
The tree tool skipped to the next entry right after printing a directory,
and only printed abstracts for files, but the storage layer only fills
abstracts for directories (files always come back empty). The flag
therefore never printed anything. Print the abstract after either kind
of entry and ask for up to 1024 characters, enough for a full skill
description, so tree(uri="viking://~/skills", level_limit=1,
include_abstract=true) lists every skill with its description.
Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28
* fix(skills): honor node_limit in GET /api/v1/skills
list_skills declared node_limit but always listed each skill root with a
hardcoded 1000. Pass it through per root; 0 keeps the default so the CLI's
accepted range (-n 0) still lists everything.
Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28
* fix(mcp): point skill hits in find/search at their SKILL.md
A skill is indexed through its directory's .abstract.md, so find and
list-mode search printed hits like viking://agent/skills/x/.abstract.md.
Following the "use the read tool to expand a URI" advice returned only
the frontmatter, and read_content inlined the same stub. Skill hits now
show <dir>/SKILL.md, and read_content reads that file.
Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28
* fix(mcp): validate add_skill targets and sources before minting an upload
Review findings on the add_skill tool:
- target_uri passed the content-kind check for any path under a skills
root (viking://~/skills/pdf) and, for ROOT, for another user's root,
but the installer only accepts the caller's own skills root or
viking://agent/skills. On the local-path branch the tool minted a
one-time upload token anyway, and the upload failed with 400 after the
token was spent. The target is now resolved with the installer's own
rule first; shared subpaths map to viking://agent/skills, the rest fail
at once, and the error names both allowed roots.
- Non-Git remote sources such as tos:// were treated as remote, then
refused as "direct host filesystem paths". add_skill now decides Git
with the same prefixes resolve_skill_source uses (shared as
GIT_SKILL_SOURCE_PREFIXES) and reports other schemes as unsupported.
- With list_only, the upload instructions still said the skill would be
installed and that no further call was needed; they now say the upload
only lists the source's skills.
- The zip example packaged hidden files, so .git and .env files went
into the stored skill. It now excludes VCS data, .env files,
node_modules and .DS_Store, starting from a fresh archive.
- tree(include_abstract=true) printed the "abstract is not ready"
placeholder for directories that never get an abstract, such as a
skill's scripts/. Those placeholders are skipped.
Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28
* docs(mcp): say that write only refuses the user's own skills subtree
The capability reference claimed MCP write refuses every skill URI. It
refuses the user's own skills/ subtree, but under viking://agent/skills
it writes a plain file that skips skill installation. State that, and
point shared skills at add_skill as well.
Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28
* style(skills): format skill_processor.py
Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28
* fix(mcp): return one hit per skill package in find
#5045 made a skill index as a whole package, so an item-level find now
returns one hit per file inside it. Route skill-only find through
SearchService.find_skills, which keeps the best hit per package, and
resolve every skill hit to its package's SKILL.md instead of only
rewriting the two index sidecars.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
* docs(mcp): point skill changes at add_skill in the tool descriptions
The server keeps accepting write/edit under viking://agent/skills, and
forget still removes a skill directory, so the constraint lives in the
tool descriptions: add_skill is the one entry point for creating and
updating a skill, and removal goes through ov skills remove or Studio.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
* fix(mcp): describe a skill hit by its own abstract
A package hit can be any file inside the skill, whose abstract describes
that file and not the skill, so find would list a skill under a helper
script's summary. Read the package's abstract for those hits, the way
GET /skills/find already does. Keep a filter-only skill query on the
generic find, which find_skills does not serve.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
* docs(mcp): document package-level skill retrieval in find
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
* fix(mcp): do not paste an unready abstract over a skill hit's own summary
fs.abstract returns a placeholder string rather than raising when a package
has no usable .abstract.md, so the substitution replaced a useful file
summary with a diagnostic line. Reject the same placeholders tree already
rejects, and bound the per-package reads the way read_content is bounded.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
* docs(mcp): correct how search reports skill hits, and refresh the tool table
find and search both render one line per skill package, so the earlier
wording — that search returns several hits per package — contradicted the
code. Say what actually differs: search still spends a limit slot per
matching file and keeps that file's summary. Also point forget and
add_resource at add_skill where an agent would look for them, name the REST
delete alongside the CLI, and bring the capability table's line citations
back in step with mcp_endpoint.py.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
* refactor(mcp): drop guards and prose no caller can reach
install_skills only ever returns a dict, add_skill always fills root_uri,
and a source with no SKILL.md raises before it gets here, so the
isinstance, empty-list and missing-uri branches were unreachable.
fs.abstract only returns the directory placeholder. One skill package
resolves to one rendered item, so the pending map holds one each. In the
docstrings, drop what Args already says and the one removal path an MCP
caller cannot take.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
* docs(mcp): say what list-mode search actually reports for a skill hit
The search row claimed a skill package's summary is the matching file's.
It is not: _format_search_result rewrites every skill hit onto the
package's SKILL.md, keeps the best-scored one per package, and
_describe_skills_by_package replaces the summary with the package's own
abstract. What is true is that limit applies during retrieval, before
that merge, so a package matching several files still spends several
slots and fewer than limit results come back.
Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
* fix(storage): stop hiding user directories named tasks or _system from ls/tree/glob
STORAGE_INTERNAL_ENTRY_NAMES was applied as a name blacklist at every
level below the account root, so a user directory literally called
"tasks" or "_system" (e.g. viking://resources/tasks) could be created,
written, stat'ed and searched but never appeared in ls, tree or glob.
The account root already uses the LISTABLE_SCOPES whitelist, and the
internal task store lives under /local/{account}/_system/tasks, so the
only entries that must stay hidden below the root are the multi-write
lock/redirect/sync-log files. Restrict the blacklist to those.
* style: ruff format
* fix(storage): reject user writes to multi-write internal file names
.path.ovlock, .exact.ovlock.*, .redirect.json and .sync_log.json are
RAGFS multi-write metadata that listings hide at every level. The
content write path only rejected .path.ovlock, and only via the create
extension whitelist; .redirect.json could be created by a user and was
routed to the raw backend. Hidden names must not be writable.
* fix(storage): one predicate for multi-write internal names across ls/glob/write/mkdir/cp/mv/webdav
Replace STORAGE_INTERNAL_ENTRY_NAMES with is_storage_internal_name(),
mirroring RAGFS is_hidden_internal_name: .path.ovlock, .exact.ovlock.*,
.redirect.json, .sync_log.json. Listing (ls/tree/glob) now also hides
the exact-lock prefix, and the same names are rejected as mkdir/cp/mv
targets and in WebDAV paths, not only in content write.
* fix(storage): use is_storage_internal_name in resource_diff after rebase
main (#5175) added resource_diff._is_excluded_rel_path on the removed
STORAGE_INTERNAL_ENTRY_NAMES constant. Switch it to the shared predicate:
multi-write metadata stays excluded from snapshots, user files under a
directory named tasks or _system are business content.
* fix(storage): reject reserved names in ancestor directories
Support standalone and distributed vector storage with canonical index options, safe SQL translation, durable metadata, and catalog-verified ANN indexes. Separate backend responsibilities, reject unsupported distributed quantization, and follow the current deletion contract without an implicit row cap.
* fix(config): warn when unknown config field are ignored.
* test(config): fix warning capture and trim duplicate assertions
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* chore: remove session used reporting
* docs: remove stale 'used' from architecture diagram
The session/used reporting API was removed in this branch, so drop the
leftover `add/used` label from the Session box in both en/zh architecture
diagrams.
* test: drop endpoint-removal guard test
The dedicated test only asserted that POST /sessions/{session_id}/used
returns 404 after removal, which adds little value now that the endpoint
and its handler are gone. Backward-compat coverage for legacy queued
messages (usage_uris) is kept in test_session_commit_resume.py.
* retrieval: add Jev rerank provider
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* retrieval: log Jev rerank payloads
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* retrieval: make rerank payload logging configurable
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* retrieval: support Jev through Vercel gateway
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* retrieval: route Jev through Vercel's TypeSafe-compatible endpoint (#5256)
Vercel AI Gateway exposes https://ai-gateway.vercel.sh/typesafe, which
accepts TypeSafe's own System One request/response shapes. Drop the
Vercel-specific branch (undocumented /v4/ai/evaluation-model path and
SDK-internal ai-gateway-* headers) so the adapter speaks one protocol;
Vercel is now just a different api_base and model id.
Auto-detect: any api_base containing "typesafe" resolves to jev.
Docs: point Vercel config at the /typesafe base and note the long-lived
API key and credit-card requirements.
Verified live against Vercel with a real gateway key.
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: Zayn Jarvis <zaynjarvis@gmail.com>
hydrate_incremental_records filtered on the `id` field via In("id", ...),
but `id` is the collection primary key and is intentionally not part of the
ScalarIndex. VikingDB only accepts `must` filters on ScalarIndex fields, so
strict backends reject the query with:
Invalid filter param: field 'id' does not support op 'must'
This made the add_resources incremental path fail whenever the target already
had vector records to hydrate (the failing scalar query raised before the
by-id fetch fallback could run, so the fallback was effectively dead code).
Fetch the records by primary key directly (`_strict_transfer_get`) instead of
filtering on `id`: no scalar index needed, no ordering, no offset pagination.
The existing client-side projection still drops vector/sparse_vector/content,
so the returned rows and the plan snapshot are unchanged.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* fix(studio): distinguish task processing time from waiting time
* fix(tasks): exclude scheduler and downstream waits from processing time
* fix(studio): group task durations in a single column
* fix(tasks): exclude circuit breaker cooldown from processing time
Short-lived sessions created by agent integrations cannot supply a policy at creation time, so deployments need one validated fallback that is persisted into each new session. Keep that fallback separate from persisted UserConfig until the user-settings surface is complete.
Constraint: Persisted user policy and Admin API wiring remain pending maintainer scope confirmation
Rejected: Add auto_commit_policy directly to UserConfig | it would silently expose an incomplete persisted-user feature
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Explicit session policy, including null disable, must stay above the deployment fallback; idle_enabled only gates the idle scanner
Tested: 12 focused policy tests, real config-loader regression, Ruff, diff-check, and two independent review rounds
Not-tested: Native ASGI session suite because the local PersistStore extension is not built; persisted-user and doctor follow-ups
Co-authored-by: czyyyy <255856754+kwistzzqq-byte@users.noreply.github.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>