Files
OpenViking/examples/pi-coding-agent-extension/README.md
T
2cc96e393e feat(retrieval): assemble auto-recall context server-side via /search mode="context" (#3534)
* feat(retrieval): assemble auto-recall context server-side via /search mode="context"

Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.

This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.

New assembly kernel under openviking/retrieve/context_assembler/:

- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
  readable floor, then overview, then full for high-scoring entries. An
  oversized tier falls back to the previous one instead of being truncated,
  bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
  Summary section, code files reuse code_outline signatures, long documents use
  a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
  directories carry no stored abstract; their full tier stays capped at
  overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
  presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
  every harness inherits cross-turn dedup; exclude_uris remains as the
  stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
  element per entry. Every tier carries its URI, so the model can always drill
  down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
  timeout fuses, and a failed rewrite still returns the unrewritten block.
  Retrieval failures are counted into stats rather than silently yielding an
  empty block.

Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.

* refactor(retrieval): give context tiers a per-category default

The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.

Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.

Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".

Assembly fixes found alongside:

- Removing the abstract cap exemption would turn an oversized abstract into
  a bare URI, so it now falls back to overview first — for memory that is a
  cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
  `asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
  such attribute; usage was structurally always null. It now reads the model
  instance's tracker and reports only when the call count moved by exactly
  one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
  fail, and the file was never rewritten, so it could not heal. Records are
  now coerced on read and dropped on the next write, along with records left
  ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
  to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
  forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
  `viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
  bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
  started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
  resolved a different profile than `POST /recall`; an unknown `detail` value
  raised `KeyError` through the whole call instead of degrading.

* feat(codex): inject profile context on session start

Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(retrieval): raise rewrite timeout default to 30s

* docs(agents): document low-latency recall settings

* fix(codex): prefer luna as recall compressor fallback

* refactor(plugins): unify recall compression setting

* feat(plugins): enable recall compression by default

* docs(agents): use absolute links in image docs

* fix(retrieval): address context assembly review feedback

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* test: trim redundant context assembly coverage

* fix(retrieval): address second-round context assembly review

- Drop the backticked `/search` from the deprecated-recall row in both API
  overviews. The reference checker scans the whole row after the method cell
  for backticked paths, so it read the description as a route named
  `POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
  belongs to the Rust CLI, which writes `root_api_key`, `output`,
  `echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
  two Python readers had drifted into stricter subsets, so the shipped example
  already failed to load in both. Adding the new `plugin` section to a working
  ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
  Retrieval validates query and image_url before searching, and the gather
  fuse swallowed that rejection along with genuine scope failures, so a body
  of `{"mode":"context"}` came back 200 with an empty block instead of the
  documented parameter error. Runtime failures still degrade into
  `stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
  server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
  than the 30s fuse, so a rewrite that finished inside its own budget was
  aborted client-side, discarding the whole response — including the
  uncompressed block the server returns when a rewrite fails — and falling
  back to `/recall`. The deadline is only extended when the body actually
  requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
  `plugin.recallContextTimeoutMs` pins it.

* chore(plugins): sync shared modules into the zcode snapshot

* fix(retrieval): align context quotas and plugin defaults

Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(retrieval): preserve recall compatibility

Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-05 12:31:21 +08:00

19 KiB
Raw Permalink Blame History

OpenViking Memory Extension for Pi Coding Agent

Long-term semantic memory and context takeover for pi sessions, powered by OpenViking. Recall happens automatically before every prompt, capture happens after every turn, and OpenViking can own long-term context by replacing committed history with an archive overview in pi's context hook.

Design informed by lessons from all three OpenViking agent plugins: synchronous recall from OpenClaw, production-hardened capture/ranking from Claude Code, and anti-patterns dodged from Hermes's stale prefetch approach. See DESIGN.md for the base design and TAKEOVER.md for the context-takeover layer.

Quick Start

Prerequisites

  • pi coding agent installed (npm i -g @earendil-works/pi-coding-agent)
  • Node.js 18+ (for the extension's TypeScript runtime)
  • An OpenViking server reachable — local or remote

1. Have an OpenViking server reachable

Either run one locally or point at a remote one. The quickstart guide walks through both options. Default port is 1933; local mode runs without authentication.

Verify it's up:

curl http://localhost:1933/health   # or your remote URL

2. Install the extension

Use the shared installer:

bash examples/memory-plugin-shared/install.sh --harness pi

The installer copies the extension to ~/.pi/agent/extensions/openviking and registers it with pi install. The extension loads on next pi invocation.

3. Configure (optional)

Credentials are resolved from OPENVIKING_* environment variables, ~/.openviking/ovcli.conf, then ~/.openviking/ov.conf. Run the setup wizard when you need to configure a remote server:

node ~/.pi/agent/extensions/openviking/scripts/setup.mjs

~/.pi/agent/extensions/openviking/config.json is for behavior knobs only:

{
  "enabled": true,
  "syncTurns": true,
  "recallTokenBudget": 2000,
  "scoreThreshold": 0.35,
  "minQueryLength": 3,
  "profileTokenBudget": 10000,
  "resumeContextBudget": 32000,
  "commitTokenThreshold": 20000,
  "takeover": {
    "enabled": true,
    "tokenThreshold": 30000,
    "keepRecentTurns": 3,
    "overviewBudget": 3000,
    "overviewPollMs": 2000,
    "overviewPollMax": 15
  }
}

Credential environment variables:

Env Var Meaning
OPENVIKING_URL OpenViking server URL
OPENVIKING_API_KEY / OPENVIKING_BEARER_TOKEN Bearer token
OPENVIKING_ACCOUNT Trusted-mode account
OPENVIKING_USER Trusted-mode user
OPENVIKING_PEER_ID Actor peer id
OPENVIKING_WORKSPACE_PEER Derive an actor peer from the current workspace by default; set 0 to disable
OPENVIKING_RECALL_PEER_SCOPE all recalls other project memories with a score penalty; actor only sees global plus the current project

Recall asks the server to assemble the context block in one request (POST /api/v1/search/search with mode="context"), so token budgeting, detail tiers and cross-turn dedup are shared with every other harness. Deployments without that endpoint fall back to /api/v1/search/recall, and that outcome is cached so only the first turn pays for the probe.

API keys are sent as Authorization: Bearer .... By default the extension derives a peer from the process workspace path using Claude's project-directory naming rule: every non-letter-or-digit character becomes -, with no path normalization. For example, /Users/x/Dev/OpenViking becomes -Users-x-Dev-OpenViking. The effective peer is sent as X-OpenViking-Actor-Peer and stored as peer_id on captured session messages. OPENVIKING_PEER_ID overrides the workspace-derived value.

Recall defaults to the broad mode: global memory, the current workspace, and other workspace memories can all be recalled, with other workspaces penalized and rendered later. Set OPENVIKING_RECALL_PEER_SCOPE=actor for the isolation mode, which only sees global memory plus the current workspace. In deployments where one bot serves multiple real people, such as zouk, vikingbot, or AstrBot, use the isolation mode with an explicit actor peer so one person's memories are not recalled into another person's session.

4. Start Pi

pi

The extension shows an [OpenViking] status line on startup. Tools (viking_search, viking_remember, etc.) are registered automatically. Memories persist across sessions — no additional setup.

Configuration Reference

Tuning fields

All fields below live in config.json. Defaults are shown.

Field Default Description
enabled true Set false to disable the extension entirely
syncTurns true Enable auto-capture of conversation turns

Recall tuning

Field Default Description
recallTokenBudget 2000 Token budget for inline recall content
recallMaxContentChars 500 Per-item content cap for search results
recallPreferAbstract true Prefer L0 abstract over L2 full body when available
recallLimit 10 Legacy quota-scaling input converted to six coding quotas, not a final cap
scoreThreshold 0.35 Min relevance score (0–1)
minQueryLength 3 Skip recall for queries shorter than N characters

Explicit recallLimit values from 1 through 5 produce an effective total quota of 6 because each coding category keeps one retrieval slot. Direct API integrations should configure category quotas when they need exact ceilings.

Capture tuning

Field Default Description
captureMode "semantic" "semantic" (always capture) or "keyword" (trigger-based)
captureMaxLength 24000 Max sanitized text length for the capture decision
captureAssistantTurns true Include assistant turns (text + tool USE inputs)
captureToolResults false Include tool result output (noisy — off by default)
captureToolMaxChars 2000 Max captured output chars for one tool part
commitTokenThreshold 20000 Pending-token threshold for client-driven commit
commitKeepRecentCount 10 Live tail kept after commit

Context takeover

Takeover is enabled by default. OpenViking commits archived history, polls the session overview, then the context hook replaces covered conversation turns with a synthetic [OpenViking Session Context] user message while keeping the recent live tail.

Field Default Description
takeover.enabled true Let OpenViking own long-term context through the context hook
takeover.tokenThreshold 30000 Synced-token pressure that triggers commit and boundary advance
takeover.keepRecentTurns 3 Recent user turns retained in full fidelity
takeover.overviewBudget 3000 Token budget for the injected archive overview
takeover.overviewPollMs 2000 Delay between overview polling attempts after commit
takeover.overviewPollMax 15 Max overview polling attempts before fail-open

Injection tuning

Field Default Description
profileTokenBudget 10000 Token budget for user profile block
resumeContextBudget 32000 Token budget for archive overview on session resume

Misc

Field Default Description
bypassPatterns [] Glob patterns to skip extension processing
logLevel "error" "silent", "error", or "info"

Architecture

┌──────────────────────────────────────────────────────┐
│                    Pi Coding Agent                    │
│                                                      │
│  session_start  before_agent_start  context  turn_end│
│  session_before_compact  session_shutdown            │
└────────┬──────────────────┬───────────┬──────────────┘
         │                  │           │
         │  ┌───────────────▼───────────▼────────┐
         │  │   extension modules (.ts)           │
         │  │   client / sync / recall / tools    │──────►  OpenViking
         │  └─────────────────────────────────────┘        Server
         │                                                (HTTP API)
         │  ┌──────────────────────────────────────┐
         └──►  7 registered LLM tools              │
            │  viking_search / viking_read / …     │
            └──────────────────────────────────────┘

The extension is a single directory of TypeScript files loaded by pi's jiti transpiler — no build step, no npm dependencies, no MCP server. All communication goes over HTTP to the OpenViking REST API.

Event Flow

Pi Event Extension Action
session_start Health check → derive OV session → build profile context → restore takeover state
before_agent_start Idempotent startup for pi -c + queue the current prompt for recall
context Run current-prompt recall after UI rendering, then inject takeover and recall context
turn_end Extract branch entries → write or pending-queue OV messages → maybe advance boundary
session_before_compact Takeover mode returns OV overview as pi compaction summary; otherwise commits pending messages
session_shutdown Persist takeover state or final non-takeover commit

Recall: Synchronous, Not Stale

Unlike Hermes's stale prefetch (recall from previous turn's query, injected one turn late), this extension searches OpenViking with the current user prompt via pi's context event. Pi renders the submitted user message before this hook, so recall latency does not hold the message off-screen. Results are still injected into the same model turn as <openviking-context> blocks. This means:

  • First turn of a session gets relevant context immediately
  • Topic switches within a session get correct recall
  • No waiting for the next turn to see relevant memories

Memory Pollution Prevention

Before pushing turns to OpenViking, shared capture sanitization strips injected context blocks such as <openviking-context> to prevent a self-referential pollution loop where recall context is captured back as user messages.

In takeover mode the adapter uses faithful capture: acknowledgments and short turns are retained because they may later be represented only through the OV archive overview. Empty text, slash commands, and OpenViking status messages remain filtered.

Tool Use Preservation

Tool capture preserves structured tool parts with bounded inputs and outputs. The memory extractor sees what the agent did without indexing unbounded raw output.

LLM Tools

The extension registers 7 tools that pi's model can invoke on demand:

Tool Description
viking_search Semantic search across memories, resources, and skills
viking_read Read a viking:// URI at abstract / overview / full level
viking_browse List directory contents or stat a viking:// URI
viking_remember Store a fact or preference into long-term memory
viking_forget Delete a memory by URI or search query
viking_add_resource Ingest a URL into OpenViking for indexed retrieval
viking_archive_expand Expand an archived session back into raw conversation

The canonical /viking command (type /viking in pi's chat) displays connection status, session info, and accepts commit for manual synchronous commit.

Compared to Pi's Built-in Memory

Pi has a built-in MEMORY.md file system. This extension complements it:

Feature Built-in MEMORY.md OpenViking extension
Storage Flat markdown Vector DB + structured extraction
Search Loaded into context wholesale Semantic similarity + ranking + token budget
Scope Per-project Cross-project, cross-session, cross-agent
Capacity Context-limited Unlimited (server-side storage)
Extraction Manual rules LLM-powered entity / preference / event extraction
Subagents Same as parent Isolated session + typed agent namespace

Compared to Claude Code Plugin

Both plugins share the same core design (informed by each other):

Feature Claude Code Plugin Pi Extension
Architecture Hook scripts (.mjs) + MCP delegation Native TypeScript extension
Recall timing Synchronous (UserPromptSubmit hook) Synchronous (context event)
Tool delivery OV server's MCP endpoint (16 tools) pi.registerTool() (7 tools)
Write path Detached worker (async) Async promise (pi's event loop)
Installation claude plugin install + setup script Copy directory → auto-discovered
Memory index None (flashlight search model) Built (map model — model sees what OV knows)
Subagent isolation Explicit hook management Natural process-level isolation

Extension Structure

See DESIGN.md for the full design specification — comparison of all three OV plugins, detailed event flow, design rationale, and implementation guidance useful for building OV extensions for any agent harness.

pi-coding-agent-extension/
├── config.json          # Default configuration (edit to customize)
├── config.ts            # Config loader (defaults + config.json merge)
├── client.ts            # OpenViking HTTP client (fetch + response envelope)
├── sync.ts              # Turn capture, write queue, session lifecycle
├── recall.ts            # Synchronous recall with ranking + budget
├── takeover.ts          # Thin pi binding around lib/takeover-core.mjs
├── tools.ts             # 7 registered LLM tools + /viking command
├── lib/takeover-core.mjs # Pure context-takeover state machine
├── index.ts             # Extension entry point (event handlers)
├── TAKEOVER.md          # Context-takeover design
└── README.md

All TypeScript files are loaded directly by pi's built-in jiti transpiler — zero dependencies beyond Node.js.

Troubleshooting

Symptom Cause Fix
Extension not loading enabled: false in config.json Set "enabled": true
No recall on first prompt OpenViking server not running or wrong URL curl http://localhost:1933/health
Tools not showing after pi -c resume Known pi issue (tools not re-registered on resume) Workaround built in — tools register in before_agent_start
Extension crashes on load Wrong OV server URL or network issue Check logLevel and server accessibility
No memories extracted Wrong embedding/extraction model in OV config Check OV's embedding / vlm configuration
Takeover never advances Pending addMessage replay, commit, or overview polling failed Set OV_DEBUG_LOG=/tmp/ov-pi.log and retry /viking commit

License

Apache-2.0 — same as OpenViking.