Files
OpenViking/examples/common/recipe.py
T
9eac8a6d3d feat(sdk): sync go/ts/python SDKs with server find/search, recall, an… (#3737)
* feat(sdk): sync go/ts/python SDKs with server find/search, recall, and admin changes

Server-side changes recently landed that the language SDKs had drifted from:

- find/search results now return `tags` and no longer return
  `category`/`match_reason`/`relations`/`overview` (#3730). Go's strict
  struct was the only one broken; update MatchedContext accordingly.
- new admin endpoints for agent-evolution and per-account settings (#3695).
- public `search/recall` endpoint was missing from all SDKs.

Changes:
- python: add `level`/`since`/`until`/`time_field` to find/search; add an
  `extra` escape hatch to find/search/add_resource/write/batch_write so new
  server fields can be passed without an SDK bump (only forwarded when set,
  preserving `level=0`); add `recall` and the four admin methods.
- go: fix MatchedContext (add Tags, drop removed fields), add Recall and the
  four admin methods.
- typescript: type MatchedContext/FindResult, add RecallOptions, add `recall`
  and the four admin methods.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): unify options APIs and sync latest server interfaces

- migrate complex Python SDK calls to typed options dictionaries
- add dedicated context search and consistent extra-field handling
- align Go and TypeScript options with omission-aware serialization
- support session config, event tags, Agent Evolution date filters,
  OpenViking Assets, batch write, downloads, and create_parent
- refresh SDK tests and examples across all three languages

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): address options API review findings

- fix Go session extra merging and Python message precedence
- adapt LangChain calls to the Python options API
- migrate repository examples, tests, and documentation

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): complete options migration and message parity

- migrate remaining Python SDK benchmarks to options dictionaries
- normalize empty parts consistently for single and batch messages
- add regression guards for repository SDK call sites

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): align reindex options after main rebase

- preserve reindex tags in Python typed options
- add reindex extra support for Go and TypeScript
- reject official fields passed through extra across SDKs

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): support legacy keyword options

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

docs(sdk): use explicit Python SDK arguments

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): support set tags extra options

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): expose Go add resource options

Expose AddType and ProcessingMode through Go AddResourceOptions and serialize them to the resources API. Add a regression test covering the resulting request payload.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): flatten core Python client options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

docs(sdk): align Python examples with flattened options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): preserve core API compatibility

Co-authored-by: TRAE CLI <traecli@bytedance.com>

refactor(python-sdk): move resource hints to options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): align resource option callers

Co-authored-by: TRAE CLI <traecli@bytedance.com>

test(sdk): cover recursive reindex forwarding

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): preserve Go options compatibility

Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(python-sdk): expose message peer id

Co-authored-by: TRAE CLI <traecli@bytedance.com>

test(python-sdk): consolidate options coverage

Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(python-sdk): add parts and flatten image search

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* docs(sdk): align Python call examples

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-08-24 14:09:11 +08:00

258 lines
8.6 KiB
Python

#!/usr/bin/env python3
"""
RAG Pipeline - Retrieval-Augmented Generation using OpenViking + LLM
Focused on querying and answer generation, not resource management
"""
import json
import time
from typing import Any, Dict, List, Optional
import requests
from openviking_sdk import SyncHTTPClient
class Recipe:
"""
Recipe (Boring name is RAG Pipeline)
Combines semantic search with LLM generation:
1. Search OpenViking database for relevant context
2. Send context + query to LLM
3. Return generated answer with sources
"""
def __init__(
self,
config_path: str = "./ov.conf",
server_url: str = "http://127.0.0.1:1933",
):
"""
Initialize RAG pipeline
Args:
config_path: Path to config file with LLM settings
server_url: OpenViking HTTP server URL
"""
# Load configuration
with open(config_path, "r") as f:
self.config_dict = json.load(f)
# Extract LLM config
self.vlm_config = self.config_dict.get("vlm", {})
self.api_base = self.vlm_config.get("api_base")
self.api_key = self.vlm_config.get("api_key")
self.model = self.vlm_config.get("model")
# Initialize OpenViking client
self.client = SyncHTTPClient(url=server_url)
self.client.initialize()
def search(
self,
query: str,
top_k: int = 3,
target_uri: Optional[str] = None,
score_threshold: float = 0.2,
) -> List[Dict[str, Any]]:
"""
Search for relevant content using semantic search
Args:
query: Search query
top_k: Number of results to return
target_uri: Optional specific URI to search in. If None, searches all resources.
score_threshold: Minimum relevance score for search results (default: 0.2)
Returns:
List of search results with content and scores
"""
# print(f"🔍 Searching for: '{query}'")
# Search all resources or specific target
# `find` has better performance, but not so smart
results = self.client.search(
query=query,
target_uri=target_uri or "",
options={
"score_threshold": score_threshold,
},
)
# Extract top results
search_results = []
resources = results.get("resources", [])[:top_k]
memories = results.get("memories", [])[:top_k]
for _i, resource in enumerate(resources + memories): # ignore SKILLs for mvp
uri = resource["uri"]
score = resource.get("score", 0.0)
try:
content = self.client.read(uri)
search_results.append(
{
"uri": uri,
"score": score,
"content": content,
}
)
# print(f" {i + 1}. {uri} (score: {score:.4f})")
except Exception as e:
# Handle directories - read their abstract instead
if "is a directory" in str(e):
try:
abstract = self.client.abstract(uri)
search_results.append(
{
"uri": uri,
"score": score,
"content": f"[Directory Abstract] {abstract}",
}
)
# print(f" {i + 1}. {uri} (score: {score:.4f}) [directory]")
except:
# Skip if we can't get abstract
continue
else:
# Skip other errors
continue
return search_results
def call_llm(
self, messages: List[Dict[str, str]], temperature: float = 0.7, max_tokens: int = 2048
) -> str:
"""
Call LLM API to generate response
Args:
messages: List of message dictionaries with 'role' and 'content' keys
Each message should have format: {"role": "user|assistant|system", "content": "..."}
temperature: Sampling temperature (0.0 to 1.0)
max_tokens: Maximum tokens to generate
Returns:
LLM response text
"""
url = f"{self.api_base}/chat/completions"
headers = {"Content-Type": "application/json", "Authorization": f"Bearer {self.api_key}"}
payload = {
"model": self.model,
"messages": messages,
"temperature": temperature,
"max_tokens": max_tokens,
}
print(f"🤖 Calling LLM: {self.model}")
response = requests.post(url, json=payload, headers=headers)
response.raise_for_status()
result = response.json()
answer = result["choices"][0]["message"]["content"]
return answer
def query(
self,
user_query: str,
search_top_k: int = 3,
temperature: float = 0.7,
max_tokens: int = 2048,
system_prompt: Optional[str] = None,
score_threshold: float = 0.2,
chat_history: Optional[List[Dict[str, str]]] = None,
) -> Dict[str, Any]:
"""
Full RAG pipeline: search → retrieve → generate
Args:
user_query: User's question
search_top_k: Number of search results to use as context
temperature: LLM sampling temperature
max_tokens: Maximum tokens to generate
system_prompt: Optional system prompt to prepend
score_threshold: Minimum relevance score for search results (default: 0.2)
chat_history: Optional list of previous conversation turns for multi-round chat.
Each turn should be a dict with 'role' and 'content' keys.
Example: [{"role": "user", "content": "previous question"},
{"role": "assistant", "content": "previous answer"}]
Returns:
Dictionary with answer, context, metadata, and timings
"""
# Track total time
start_total = time.perf_counter()
# Step 1: Search for relevant content (timed)
start_search = time.perf_counter()
search_results = self.search(
user_query, top_k=search_top_k, score_threshold=score_threshold
)
search_time = time.perf_counter() - start_search
# Step 2: Build context from search results
context_text = "no relevant information found, try answer based on existing knowledge."
if search_results:
context_text = (
"Answer should pivoting to the following:\n<context>\n"
+ "\n\n".join(
[
f"[Source {i + 1}] (relevance: {r['score']:.4f})\n{r['content']}"
for i, r in enumerate(search_results)
]
)
+ "\n</context>"
)
# Step 3: Build messages array for chat completion API
messages = []
# Add system message if provided
if system_prompt:
messages.append({"role": "system", "content": system_prompt})
else:
messages.append(
{
"role": "system",
"content": "Answer questions with plain text. avoid markdown special character",
}
)
# Add chat history if provided (for multi-round conversations)
if chat_history:
messages.extend(chat_history)
# Build current turn prompt with context and question
current_prompt = f"{context_text}\n"
current_prompt += f"Question: {user_query}\n\n"
# Add current user message
messages.append({"role": "user", "content": current_prompt})
# Step 4: Call LLM with messages array (timed)
start_llm = time.perf_counter()
answer = self.call_llm(messages, temperature=temperature, max_tokens=max_tokens)
llm_time = time.perf_counter() - start_llm
# Calculate total time
total_time = time.perf_counter() - start_total
# Return full result with timing data
return {
"answer": answer,
"context": search_results,
"query": user_query,
"prompt": current_prompt,
"timings": {
"search_time": search_time,
"llm_time": llm_time,
"total_time": total_time,
},
}
def close(self):
"""Clean up resources"""
self.client.close()