* feat: add freshness-aware parent aggregation
Defer wide-directory abstract/overview regeneration until the configured freshness threshold is reached while continuing changed-file semantic and vector processing.
Persist freshness metadata atomically, make parent bubbling L0-aware, preserve separate semantic/vector statuses, and keep explicit waits synchronous.
Rebuild every sampled summary on threshold refresh and always retry directory vectorization so stale sidecars or transient vector failures cannot be silently accepted.
Add focused coverage for freshness policy, pending-state consumption, sampled-summary refresh, vector retries, and parent bubbling.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive, and applied to memory
* feat: ov reindex support --recursive, and applied to memory
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* chore: remove dead git tuning knobs and duplicate release/frontend files
- GitTuningConfig: drop upload_concurrency, restore_concurrency,
ref_cas_max_retry, and ref_cas_backoff_ms, which were parsed but never read
anywhere (verified no source readers). Keep commit_index_enabled and
blob_exists_precheck_enabled, the two knobs that actually take effect.
- Design doc: mark the removed knobs as roadmap items to be added back when
the behavior lands, and correct stale 'not implemented' claims
(validate_account_id is enforced at the Git service entry; blob reads are
limited via show_with_limit).
- Remove .github/workflows/release-vikingbot-first.yml: a historical one-off
PyPI release workflow whose bot/ package lacks build metadata.
- web-studio: delete pnpm-lock.yaml and the pnpm-only package.json block;
the Makefile and CI already use npm + package-lock.json as the single
install chain. Net -9,695 lines.
* docs: align cleanup notes with current behavior
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
Use ZCode rollout logs as the authoritative incremental source, advance capture state only for the acknowledged prefix, and persist host turn identity with the OpenViking turn_id contract.
Detach Stop writes, package ZCode in the TOS marketplace artifact, add end-to-end regressions, and move the integration docs under community plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(integrations): add ZCode memory plugin
Add examples/zcode-memory-plugin — a thin ZCode lifecycle adapter that reuses
the shared memory-plugin-shared runtime for recall, capture, commit, and MCP
proxy. No memory logic is duplicated.
Key design decisions (see docs/design/zcode-memory-plugin-design.md):
- Vendor shared runtime into scripts/shared/ via sync.mjs (self-contained plugin)
- 4 hook events only (SessionStart, UserPromptSubmit, PreToolUse, Stop) —
ZCode does not support PreCompact/SessionEnd/SubagentStart/SubagentStop
- Output schema: ZCode-canonical keys only (no Claude-Code 'decision: approve')
- Config-driven install: hooks + MCP merged into ~/.zcode/cli/config.json
- install.sh wiring: detection, TUI, validation, install, uninstall
Verified locally:
- 22/22 node:test cases pass (turns parser + hook output schema)
- sync.test.mjs passes
- install → uninstall cycle: hooks/MCP correctly written and cleaned
- URI guard denies viking:// paths with MCP redirect
- Capture writes to OV session with zc- prefix
Closes#3127
Related: #3442, #3544
* chore: remove non-essential files from PR, add .scratch to .gitignore
- Remove .scratch/ working notes (local ticket files, not codebase artifacts)
- Remove package.json and .gitignore from plugin dir (TRAE/Cursor don't have them)
- Add .scratch/ to root .gitignore
* fix(zcode): use verified ZCode field names + rollout fallback for capture
- Update zcode-turns.mjs to probe responseText/responsePreview (verified
from ZCode reverse-engineering in #3127 by @quinn-zenith) instead of
the TRAE-inferred last_assistant_message
- Add rollout file fallback: when stdin payload lacks user content (the
known ZCode limitation), read ~/.zcode/cli/rollout/model-io-sess-*.jsonl
to extract the last user+assistant pair from request.messages+response
- Fix concurrent session isolation: normalize sessionId→session_id in
zcode-hook.mjs before resolveNativeSessionId to prevent cwd-fallback
collision when two ZCode windows run in the same directory
- Add 2 new test cases for rollout fallback (14 turns tests total, 24 total)
- All 24 tests pass
* fix(zcode): address maintainer review blockers (config safety, MCP ownership, turnId)
Addresses 3 blockers from @huangruiteng's review (CHANGES_REQUESTED):
1. Config safety: distinguish ENOENT from parse errors — malformed
config.json now aborts instead of overwriting. Use backup+tmp+rename
for atomic writes.
2. MCP ownership: only replace/delete mcp.servers.openviking entries
tagged as openviking-memory. User-managed entries with the same name
are preserved on install and untouched on uninstall.
3. TurnId-based dedup: rollout entries carry monotonic turnId — now used
as the primary dedup key (capturedTurnIds set) instead of stableHash.
extractUnseenRolloutTurns scans ALL unseen entries since lastTurnId,
not just the last row — recovers missed turns after hook failure.
Fail-closed when no turns are found.
Also updates DESIGN.md to reflect verified field names (responseText/
responsePreview) and the turnId contract.
27/27 tests pass (was 24). Added 3 new rollout tests: incremental
capture with lastTurnId, multi-entry scan, turnId propagation.
* docs(zcode): update stale field name references in design spec
Update test case descriptions to match verified field names
(responseText/responsePreview instead of last_assistant_message)
and add rollout fallback + turnId test coverage descriptions.
* fix(zcode): dedup key includes role + first-capture returns all turns
Fix two bugs found in code review pass 2:
1. Assistant turns silently dropped: user and assistant from the same
rollout entry shared a turnId, so dedup via capturedTurnIds dropped
the assistant. Fix: dedup key is now ${turnId}:${role}, not turnId
alone. Regression test added.
2. First-capture data loss: when no lastKnownTurnId was set, only the
last rollout entry was returned, losing prior turns. Fix: first-time
capture now returns ALL entries.
Also: add backup step to config atomic write (copyFileSync before tmp+rename),
fix line width in zcode-turns.mjs, add 2 lifecycle tests (missed Stop
recovery, user+assistant same turnId).
29/29 tests pass (was 27).
* test(zcode): add concurrent session isolation tests
Two new test cases addressing maintainer criterion 4 (concurrent sessions):
1. Two sessions read their own rollout files — verifies session A cannot
see session B's content and vice versa (sentinel-based assertion)
2. Independent lastTurnId state per session — verifies incremental capture
progresses independently when one session has prior state and another
is fresh
31/31 tests pass (was 29).
* fix(zcode): correct rollout file path pattern (model-io-<sessionId>)
The rollout path used model-io-sess-${sessionId} but ZCode filenames are
model-io-<sessionId> where sessionId already includes the sess_ prefix.
This caused the rollout fallback to always miss the file and return empty,
defeating capture entirely in production.
Verified on live two-session ZCode setup:
- Session A (sess_8c6ce483): 2 messages, 2 commits
- Session B (sess_74759710): 2 messages, 2 commits
- No cross-contamination between sessions
31/31 tests pass. Updated all test rollout filename patterns.
* docs(zcode): fix stale rollout path in comments and DESIGN.md
Comments referenced model-io-sess-<sessionId> but actual pattern is
model-io-<sessionId> (fixed in code already, comments were stale).
---------
Co-authored-by: woshiguanxiaoliang <woshiguanxiaoliang@noreply.gitcode.com>
* feat(agent-evolution): reload global switch at commit time
* feat(agent-evolution): expose configured account in status
* test(agent-evolution): cover account in status response
* fix(agent-evolution): align live config reload semantics
* fix(agent-evolution): tolerate non-object live config
* fix(usage-reporter): use snake case count fields
* feat(usage-reporter): add file log sink
* fix
* fix: address live reload and usage sink review findings
* fix(usage-reporter): complete file sink compatibility
* fix: make experience snapshot source unambiguous
* docs(usage-reporter): align count record implementation plan
* fix: address agent evolution review blockers
* fix(usage-reporter): preserve Windows rollover deadline
* fix(usage-reporter): encode file records as JSON envelopes
* fix(usage-reporter): use snake case unique id
* refactor(parse): simplify resource ingestion routing
Freeze resolved resource types before parser selection and remove unused parser extension paths so ingestion follows one documented route.
* fix(feishu): preserve sheet and bitable imports
Move Feishu-specific conversion into the accessor so the parser routing refactor keeps all supported resource types.
* fix(parse): keep normalized Feishu content internal
Prevent Feishu Markdown produced by the accessor from being sent through Understanding a second time.
* fix(feishu): parse bitable blocks embedded in sheets
Use spreadsheet metadata blockInfo instead of treating zero-sized Bitable blocks as empty sheets.
* fix(feishu): download bitable attachment images
* refactor(parse): remove unused document converter
* refactor(parse): unify Understanding routing
* docs(parse): mark routing classification points
* docs(parse): complete wait routing flow
* fix(parse): preserve Feishu Base URL scope
* refactor(resource): separate ingestion submission from execution
* fix(resource): reject internal ingestion fields at public entry
* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest
Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.
- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit
Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.
* fix(vectordb): reject stale index replacements
* fix(vectordb): harden bulk rebuild lifecycle
---------
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
The push-OTP feature (mint an OTP in Studio to hand to an MCP client) was never
wired to a consumer: consume_otp had zero production callers and no endpoint or
grant ever redeemed an OTP. The 'full happy path' test actually exercised the
display_code flow, not OTP. So the sidebar footer's 'OAuth setup' entry minted a
code with nowhere to use it — dead, confusing UX.
Remove it end-to-end and repurpose the footer slot into an entry for the
cross-device verify page (enter the 6-char display_code), which previously had no
discoverable entry point in Studio.
Frontend:
- delete oauth-setup-dialog.tsx + /oauth/setup route (+ routeTree, i18n)
- extract CrossDeviceVerifyForm from verify.tsx; add CrossDeviceVerifyDialog
- footer 'OAuth verify' entry opens the verify dialog (desktop) / page (mobile)
Backend:
- drop issue_otp route + OTPRequest/OTPResponse, storage insert_otp/consume_otp,
oauth_config.otp_ttl_seconds, and the OTP-specific tests
- keep otp.py generate_otp (cross-device display_code) + hash_secret, the shared
_atomic_consume_code, and the oauth_codes.kind column
- convert the race/expiry/revoke/GC storage tests to auth-code rows
Docs: update 11-oauth, 06-mcp-integration, and the design doc to reflect removal.
#2160 dropped the legacy `/console` standalone service but deliberately
left the OAuth authorize page's `/console` link and Quick-authorize panel
in place, calling out a follow-up to re-point them at web-studio. This
PR is that follow-up.
Backend
- `provider.authorize()` now defaults to redirecting to
`/studio/oauth/consent` (same-origin SPA) instead of the server-rendered
`/oauth/authorize/page`. New `FALLBACK_AUTHORIZE_PAGE` constant exposed
for callers that need to opt into the legacy path.
- New public endpoint `GET /api/v1/auth/oauth/pending/{pending_id}` returns
the minimum info the consent UI needs (client_name, redirect_host,
scopes); deliberately does NOT expose display_code or full redirect_uri.
- `POST /api/v1/auth/oauth-verify` now accepts either `pending_id`
(Studio consent path) or `code` (cross-device fallback).
- HTML `/oauth/authorize/page` template stripped of `/console` link, the
`/console/api/v1/...` JS, and the Quick-authorize same-origin panel.
It now serves as a pure cross-device fallback that points users at
`/studio/oauth/verify` on another already-signed-in device.
Web Studio
- New `<IdentityPicker>` shared component: "current identity" or
"use a different API key" — the temporary key is never persisted.
- New routes `/studio/oauth/consent` (same-device consent card) and
`/studio/oauth/verify` (cross-device code entry).
- ConnectionDialog gains an "OAuth client OTP" section (same
IdentityPicker), driving `POST /api/v1/auth/otp`.
- API key storage is unchanged: only sessionStorage. No new localStorage
writes, no cross-tab channels — the consent UI runs inside Studio's own
tab, so it reads the session-stored key directly.
Docs
- 11-oauth.md (zh/en): refreshed quickstart, How-it-works, Claude.ai
walkthrough, curl example, and troubleshooting around the Studio
consent / cross-device verify split.
- 12-public-access.md (zh/en): rewritten to lead with public HTTPS;
the `:1934` Caddy block is now a one-paragraph compatibility note for
deployments that already bookmarked it.
- mcp-oauth2-1.md: top-level "Studio migration" note explains the new
default path; Phase 1 history retained.
- Caddyfile / docker-compose.yml comments reworded from "aggregated
proxy" to "legacy fallback" to match the new docs.
Tests
- `tests/server/oauth/test_router.py` fixture pins to
FALLBACK_AUTHORIZE_PAGE so existing end-to-end assertions keep working.
- 4 new tests cover the pending-info endpoint and pending_id verify path.
- 55 passed locally; ruff format+check, web-studio tsc/eslint/prettier
all clean.
Security notes
- Consent UI requires explicit user click; client_name + redirect_host
shown for phishing identification.
- Knowing a pending_id does not bypass Bearer auth.
- display_code is not returned by GET pending — the cross-device
brute-force protection is preserved.
- `ctx.from_oauth` gate (router.py) untouched: OAuth bearer still
cannot mint new OAuth state or OTPs.
Replace openclaw-integration.md and openclaw-context-engine-refactor.md
with openclaw-plugin-design.md covering the three main chains (assemble,
afterTurn, compact), session mapping, tool registry, and config reference.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(oauth): hand-sewn OAuth 2.1 M1+M2 (config, JWT, storage, /oauth/token, JWT discriminator)
Snapshot before evaluating migration to mcp.server.auth SDK provider. The
hand-rolled HS256 JWT implementation in openviking/server/oauth/jwt.py is
the main candidate for replacement: its surface area is small but it would
require careful crypto review by maintainers, while the official MCP SDK
already ships an OAuth provider wired into FastMCP.
Included so far:
- OAuthConfig + integration into OpenVikingConfig (default disabled)
- openviking/server/oauth/{jwt,storage,otp,router}.py
- POST /oauth/token (authorization_code + refresh_token, PKCE S256, RFC 6749 errors)
- JWT discriminator in resolve_identity (fail-closed; ResolvedIdentity.from_oauth)
- WWW-Authenticate Bearer hint on /mcp 401 (RFC 9728)
- 49 OAuth-specific unit/integration tests (all passing)
Not yet implemented (M3 / MVP gap):
- /oauth/register (DCR), /oauth/authorize (HTML + OTP submit), well-known metadata
- POST /api/v1/auth/otp REST endpoint
* refactor(oauth): switch to mcp.server.auth SDK provider, drop hand-sewn JWT
Replaces the hand-rolled HS256 JWT signer / token endpoint / DCR with
the OAuth 2.1 surface shipped in mcp.server.auth. We supply a Provider
that adapts the existing OAuthStore (SQLite) to the SDK Protocol, plus
two custom routes the SDK doesn't own: an OTP-entry HTML page (the URL
provider.authorize() returns) and POST /api/v1/auth/otp for issuing
OTPs against an existing API key.
Net result: all OAuth crypto is now the SDK's responsibility (PKCE
S256, redirect_uri matching, error formatting). The OpenViking-side code
contains zero cryptography — access tokens are opaque random strings
prefixed with `ovat_` and looked up in SQLite by SHA-256 hash. Refresh
tokens, auth codes, OTPs use the same scheme.
Highlights:
- openviking/server/oauth/provider.py: OpenVikingOAuthProvider implements
the 8-method SDK Protocol, including subclassing AuthorizationCode /
RefreshToken / AccessToken to pin (account_id, user_id, role) per
token. Refresh-token replay triggers per-user chain revocation.
- openviking/server/oauth/storage.py: adds oauth_access_tokens and
oauth_pending_authorizations tables; peek_auth_code / peek_refresh
for non-destructive lookups; revoke_user_tokens cascades all OAuth
state for an (account, user) pair when a key is rotated.
- openviking/server/oauth/router.py: minimal authorize page (inline
HTML with frame-ancestors 'none') + OTP endpoint authenticated via
existing get_request_context dependency.
- openviking/server/auth.py: replaces JWT discriminator with prefix
match + provider.load_access_token; still fail-closed.
- openviking/server/app.py: mounts SDK routes via create_auth_routes
alongside our authorize-page + OTP routes.
- Deletes openviking/server/oauth/jwt.py and tests/server/oauth/test_jwt.py.
Tests: 32 passing, including a full DCR -> OTP -> authorize page ->
token-exchange -> /mcp lookup happy path, refresh rotation, and replay
detection. Existing test_auth.py regression unchanged.
Phase 1 still missing for full Claude.ai connectivity:
- WWW-Authenticate hint already present on /mcp 401 (from M2)
- /.well-known/oauth-protected-resource (RFC 9728) — not currently
emitted by the SDK; small custom route still TODO.
* docs(oauth): rewrite design doc to reflect mcp.server.auth SDK approach
The earlier draft described a hand-sewn HS256 JWT plan; the implementation
took a different route after discovering mcp.server.auth ships a complete
RFC 6749 / 7591 / 8414 server. Updated to reflect:
- SDK owns the protocol surface (DCR, /authorize parsing, /token, metadata,
PKCE, redirect_uri matching, error codes).
- OpenViking only contributes a Provider implementation, the OTP-entry
HTML page, and POST /api/v1/auth/otp.
- Tokens are opaque (ovat_ / ovrt_ / ovac_ prefixes) — no JWT, no crypto
on our side.
- Implementation status: M1/M2/M3 done; only RFC 9728 protected-resource
metadata + reverse-proxy issuer derivation remain for full Claude.ai
end-to-end connectivity.
* feat(oauth): add /.well-known/oauth-protected-resource (RFC 9728)
The /mcp 401 path already advertises this URL via WWW-Authenticate
Bearer resource_metadata="...", but the endpoint itself didn't exist —
clients fetched it and got a 404, which silently broke the discovery
chain even though /.well-known/oauth-authorization-server worked. Wire
up the resource metadata document so the full RFC 9728 → RFC 8414
discovery chain works end-to-end.
Uses mcp.shared.auth.ProtectedResourceMetadata pydantic model. Reads
X-Forwarded-Proto/Host so the published resource URL matches what the
client used (matches our existing WWW-Authenticate behavior).
Cache-Control: max-age=3600 — metadata is stable across requests.
* feat(console): add OTP issuance button in Settings panel
Adds a "Get OTP" button under the Settings panel of the 8020 web
console. Clicking it issues an OAuth OTP via the user's existing API
key (already loaded into sessionStorage) and displays it inline with
a copy-to-clipboard button.
Replaces the previous workflow of users having to:
curl -X POST -H "X-Api-Key: $KEY" http://1933/api/v1/auth/otp
…with a single button-click flow that the user can reach from any
machine with a browser.
Wires:
- console/app.py: new POST /console/api/v1/ov/auth/otp proxy route,
forwarding to upstream /api/v1/auth/otp. Not gated by write_enabled
since OTP issuance is an authentication artifact, not data mutation.
- index.html: new OAuth section in the Settings panel with otpBox
(hidden until OTP is generated) and a Copy button.
- app.js: getOtpBtn click handler calls callConsole, otpCopyBtn copies
to clipboard. Clear failure messages when the user has no API key
loaded yet.
This is the lightweight half of the Console-OAuth integration. The
fuller "same-origin auto-authorize" flow (Phase 2) — where the
authorize page detects sessionStorage and submits the OTP form
automatically — is still TBD and will reuse this proxy route.
* feat(oauth): device-flow style authorize page + console verify form
Pivots the OTP flow direction so the UX matches OAuth 2.0 Device
Authorization Grant (RFC 8628) more closely:
Old (push): user goes to console -> Get OTP -> copy -> paste in
client's authorize page -> submit -> redirect.
New (pull): client's authorize page DISPLAYS a 6-char code -> user
types it into the console verify form -> page polls -> redirect.
This removes one tab switch and aligns with how users mentally model
authorization ("I'm approving the request shown over there from
where I'm already signed in"). The legacy POST /api/v1/auth/otp +
"Get OTP" button are kept under a collapsed details element for any
scripted/CLI flows that still drive the older pattern.
Also wires OPENVIKING_PUBLIC_BASE_URL env var as the highest-priority
public origin override, used consistently by:
- /.well-known/oauth-protected-resource
- WWW-Authenticate header
- authorize page links
- SDK issuer at app start.
Server changes:
- storage.py: oauth_pending_authorizations gains display_code,
verified, verified_account_id/user_id/role columns; new
find_pending_by_display_code + mark_pending_verified.
- provider.authorize() now generates display_code at pending creation
and returns the page URL.
- router.py:
* GET /oauth/authorize/page — renders the code + same-origin quick-
authorize panel (sessionStorage detection, but click still required
so authorization is never silent).
* GET /oauth/authorize/page/status — polled by the page until verified;
response carries the redirect_url with auth_code on approval.
* POST /api/v1/auth/oauth-verify — authenticated; binds caller
identity to a pending row (decision=approve|deny).
Console changes:
- Settings panel: new "Authorize an MCP client" section with code input
and Authorize/Deny buttons. Legacy "Get OTP" still available under
details.
- console proxy gains POST /console/api/v1/ov/auth/oauth-verify.
Tests: 38 OAuth tests passing, including a full device-flow happy path,
deny path, idempotency (one-shot pending), unknown-code rejection,
status-410 on consumed/expired, refresh rotation, OPENVIKING_PUBLIC_BASE_URL
override, and X-Forwarded-* fallback.
* docs(oauth): add 11-oauth guide + Caddy/nginx templates + .env-driven compose
Adds a top-level OAuth 2.1 guide (zh/en) covering the production path
end-to-end. Opens with a 5-step recommended setup so readers don't have
to wade through the rationale before they can deploy. Drops the "MCP"
qualifier from the doc name — OAuth 2.1 here is generic and serves any
OAuth client, not just MCP.
- docs/{en,zh}/guides/11-oauth.md: new. Recommended setup at the top,
then background, full device flow, HTTP-local vs HTTPS-production
deployment, Caddy + nginx templates, docker-compose with the shipped
Caddy service, curl walkthrough, config reference, troubleshooting.
- docker-compose.yml: replace the prior PR's commented-out hint with a
single OPENVIKING_PUBLIC_BASE_URL var (read by both the openviking
service and an optional Caddy reverse-proxy service that's also
shipped commented-out). Same env var drives Caddy via
{$OPENVIKING_PUBLIC_BASE_URL}, so the public domain is configured
once in .env.
- docs/{en,zh}/guides/06-mcp-integration.md: replace the "OAuth Proxy
(planned, use community Cloudflare Worker)" section with a short
pointer to the new 11-oauth guide. The community proxy is still
mentioned as an alternative.
Same env-variable design also matches what the MCP add_resource tool
expects (it already reads OPENVIKING_PUBLIC_BASE_URL), so deployments
get a single source of truth for the public address.
* fix(oauth): read API key from localStorage on authorize page
The same-origin "Quick authorize" panel was reading sessionStorage,
which is per-tab. Since the OAuth authorize page opens in a different
tab from the console, the panel never showed up even when the user was
signed in.
The console persists the API key in localStorage as well (key
"ov_console_api_key" — see static/console_settings.js's
LEGACY_API_KEY_STORAGE_KEY) for cross-tab use, and that copy is what
the authorize page should consult.
Switch the page JS to localStorage first, fall back to sessionStorage
for resilience. No console-side change needed; the localStorage entry
has been written by the console all along.
* docs: add public access guide + default port 1934 aggregated proxy
- Add Caddyfile with :1934 HTTP aggregated proxy (merges 1933+8020)
- Enable Caddy service by default in docker-compose.yml on port 1934
- Add docs/{en,zh}/guides/12-public-access.md with full HTTPS setup guide
- Simplify 11-oauth.md: replace inline reverse proxy config with refs to 12
- Add HTTPS requirement callout to OAuth recommended setup
- Update 03-deployment.md to mention port 1934 as recommended entry point
* fix(oauth): address Copilot review + ruff format
- Update oauth_config.py docstrings to describe opaque tokens, not JWT
(we switched away from JWT during implementation)
- Remove unused authorize_rate_limit_per_min config field — was never
enforced anywhere in router/storage, dead config misled operators
- Wrap all OAuthStore read paths in self._lock (matching writes); the
shared sqlite3.Connection with check_same_thread=False is not safe
for concurrent cursor use across threads
- Clarify provider.exchange_refresh_token comment that replay revokes
the entire (account, user) family, not just the (client, account,
user) chain — broader blast radius is intentional
- ruff format: 8 files reformatted to satisfy CI lint
* perf(docker): add cargo + ccache cache mounts to py-builder stage
The two heavy RUN steps in py-builder (uv sync + maturin build) re-execute
on every Python source change because the upstream COPY layer for openviking/
invalidates the cache. Each rerun was ~510s + ~115s ≈ 10 min of wasted work
even though Rust/C++ source was unchanged.
Add BuildKit cache mounts so cargo and the C++ engine compilation can skip
work whose inputs are unchanged:
- Mount /cargo-target, cargo registry, and cargo git so cargo's incremental
build artifacts persist across layer reruns. Pin CARGO_TARGET_DIR so the
path stays stable when uv builds wheels in ephemeral isolated tempdirs.
- Install ccache and prepend /usr/lib/ccache to PATH so cmake (which calls
shutil.which("gcc")) resolves the ccache wrapper. ccache is path-agnostic,
so it benefits the cmake_build subdir even though setup.py recreates it
in a fresh tempdir each wheel build.
- Mount /root/.ccache so the ccache hash store persists across reruns.
Expected: hot rebuilds on Python-only changes drop step 15 from ~510s to
~60-120s (uv wheel packaging overhead remains; cargo + g++ skip on cache hit).
* perf(docker): drop redundant second maturin build step
The second RUN step in py-builder built ragfs-python a second time and
extracted its .so into the installed openviking package. This was
redundant: setup.py's build_ragfs_python_artifact() already runs maturin
during step 15 (uv sync --no-editable), and because build_meta passes
'bdist_wheel' through PEP 517, _should_require_ragfs_artifact() returns
True and the build fails closed if maturin can't produce ragfs_python.so.
The .so is then bundled into the wheel via package_data and installed
into /app/.venv on wheel install. The second step's only effect was to
overwrite the same file, costing ~115s per build.
Verified after the fact by inspecting the installed venv and importing
ragfs_python in the runtime container.
* feat(oauth): bind OAuth token lifetime to authorizing API key
Previously OAuth tokens lived independently of the API key that authorized
them. Rotating a user's key did not invalidate already-issued OAuth access /
refresh tokens, so a compromised key remained dangerous even after rotation.
Tie every OAuth token to the SHA-256 fingerprint of the API key whose holder
authorized it:
- APIKeyManager grows get_user_key_fingerprint(account_id, user_id) ->
sha256(stored_key_value). The stored value is whatever sits in
user_info["key"] (plaintext key or argon2id hash), written once on
create / regenerate and never mutated in place, so the fp is stable per
key-generation and changes the moment regenerate_key runs.
- OAuth storage gains an authorizing_key_fp column on oauth_codes,
oauth_pending_authorizations (verified_key_fp), oauth_refresh_tokens, and
oauth_access_tokens. ALTER TABLE migration guarded by PRAGMA table_info
for dev DBs that predate the field.
- Provider data classes thread the fp through authorize ->
exchange_authorization_code -> _issue_token_pair, and refresh rotation
preserves it from the consumed token's record.
- Router endpoints capture the caller's current fp at the only two
identity-binding moments: /api/v1/auth/otp (caller) and
/api/v1/auth/oauth-verify (verifier). If the manager returns None
(ROOT key, trusted-mode identity, or removed user), refuse to issue
OAuth state -- there is no key whose lifecycle we could honor.
- auth.py:_try_resolve_oauth_token recomputes the user's current fp on
every OAuth bearer auth and demands strict equality via
hmac.compare_digest. NULL / empty / mismatch all fail closed with a
401 telling the client to re-authorize.
Crypto notes: sha256 over a 256-bit-random API key (or its argon2id hash)
is preimage-safe, so an oauth.db leak does not reveal the API key. No new
secret material introduced; the fp is derived deterministically from data
that already exists.
Tests: 3 new lifecycle tests in test_auth_integration (rotation rejected,
user-removed rejected, missing-fp fail-closed), 3 new router tests
(no-fp caller / verifier rejected, fp recorded on access + refresh), 2 new
APIKeyManager tests (fp changes on rotate / vanishes on remove).
Pre-existing inserts in test_storage updated to pass _FP. 82/82 OAuth +
APIKeyManager tests pass.
* docs(oauth): document OAuth lifetime ≤ authorizing key lifetime
The fingerprint binding landed in the previous commit; users need to know
that key rotation now auto-invalidates derived OAuth tokens (no separate
revoke step) and that ROOT / trusted-mode identities cannot issue OAuth.
Updates both en and zh under docs/guides/11-oauth.md, replacing the
"operator should also revoke ..." paragraph with the new automatic
behavior + brief note on the SHA-256 fingerprint scheme.
* fix(oauth): close 4 review findings on token lifecycle
External security review of #1870 surfaced four real gaps in the OAuth
implementation. All four directly affect the lifecycle / privilege model.
P1: role downgrade did not invalidate OAuth tokens
set_role rewrites user_info["role"] without touching user_info["key"],
so the SHA-256 fingerprint binding stays valid and an ADMIN demoted to
USER continues to resolve as ADMIN. Refresh tokens keep minting fresh
ADMIN access tokens. Fixed in two places:
- auth.py:_try_resolve_oauth_token re-fetches Role.get_user_role and
rejects when the embedded role outranks the current role.
- provider.exchange_refresh_token gets a role_resolver callback (wired
in app.py to api_key_manager.get_user_role) and applies the same
gate before consuming a refresh.
Promotion remains harmless — the embedded lower privilege is still
authorized, only downgrades trigger rejection.
P1: confidential client secrets were never enforced
provider.get_client returned client_secret=None regardless of the
stored hash; the MCP SDK's ClientAuthenticator skips secret validation
when the returned client has a falsy secret, silently allowing
client_secret_basic / client_secret_post clients to authenticate with
only client_id. Real MCP clients all use "none" + PKCE per RFC 8252
§8.4 anyway, so register_client now rejects non-"none" auth methods at
DCR. Native/desktop apps can't keep secrets — PKCE is the actual
proof-of-possession.
P1: OAuth tokens could mint new OAuth grants
/api/v1/auth/otp and /api/v1/auth/oauth-verify accepted any caller
resolved through get_request_context, including identities resolved
from OAuth bearers. A stolen 1h access token could call oauth_verify
with its own pending row and walk away with a 30d refresh-token
chain — privilege time-extension. RequestContext now carries
from_oauth (mirroring ResolvedIdentity.from_oauth) and both endpoints
reject from_oauth=True with 403, forcing primary auth.
P2: GC erased refresh-token replay tombstones
gc_expired deleted "WHERE expires_at < ? OR consumed = 1" every
minute. After GC, is_refresh_known_but_consumed could not distinguish
a replay from an unknown token and exchange_refresh_token never fired
revoke_chain — defeating RFC 9700 §4.14 family revocation for late
replays. GC now keeps consumed refresh rows until their natural
expires_at; storage cost bounded by the 30d max refresh TTL.
Also adds from_oauth field to RequestContext and propagates from
ResolvedIdentity in get_request_context.
Tests: 7 new (role downgrade rejection in bearer auth + refresh path,
role promotion is harmless, confidential DCR rejected, from_oauth
rejected at OTP and oauth-verify, refresh tombstone preserved across
GC). Pre-existing test_oauth_root_can_be_used and
test_dcr_registers_client updated to match the stricter contract.
89/89 OAuth + APIKeyManager tests pass.
* fix(oauth): downgrade confidential DCR to public instead of rejecting
The previous P1.2 fix rejected DCR when token_endpoint_auth_method was
not "none", reasoning that we never enforce client_secret server-side
so accepting confidential auth methods would be a silent security
downgrade. That is the right invariant — but the rejection broke real
clients: the OAuth 2.0 default for token_endpoint_auth_method is
"client_secret_basic", and at least Claude Desktop relies on the SDK to
fill in defaults rather than explicitly setting "none". DCR for those
clients started returning 400 even though they would work fine with PKCE
(which they all use anyway).
Soft-failure design instead: accept any registered auth method, but
overwrite the stored value to "none" and log a warning. The end-state
is identical to the rejection path — every client is treated as
public+PKCE, no secret is ever stored or enforced — but Claude Desktop's
DCR no longer blows up.
Updates the test from asserting 400 to asserting that a confidential
registration is silently downgraded: stored auth_method == "none",
client_secret_hash is None.
* delete(docs): remove error file
* docs(zh): sync 03-deployment.md with English version
c5cb241f only updated docs/en/guides/03-deployment.md when introducing
the 1934 aggregated entry point and the 12-public-access.md guide. This
backfills the same changes in the Chinese version:
- Add port 1934 (Caddy aggregated entry point) to the access list
- Link to 12-public-access.md for public HTTPS setup
- Add 11-oauth.md and 12-public-access.md to "Related Documentation"
* feat(session): add account namespace policy and shared sessions
Unify namespace resolution across filesystem, indexing, and session storage.
Add account-shared session paths, role_id auth semantics, and an HTTP demo
script for the four namespace-policy combinations.
* space
* fix(pack): skip derived semantic files in ovpack transfer
Keep ovpack imports resilient to stale sidecars and rebuild semantics through the normal queue instead of restoring derived files verbatim.
* Revert "fix(pack): skip derived semantic files in ovpack transfer"
This reverts commit f4e4db8401.
* fix(namespace): default legacy accounts to agent-shared policy
Clarify that memory.agent_scope_mode is deprecated and document the supported agent memory migration paths.