chore(deps): bump ai v7, @ai-sdk/* v4, zod 4, vite 8, plugin-react 6, typescript 7 (#155)

* chore(deps): bump ai v7, @ai-sdk/* v4, zod 4, vite 8, plugin-react 6

Supersedes the six dependabot PRs for these packages. The lockfile is
rewritten wholesale: the old lock pinned vite 5 / plugin-react 4 as root
ranges, and since those two are mutual peers npm could not move them
incrementally (ERESOLVE). Regenerating resolves cleanly with no overrides.

Required migrations, all driven by upstream renames or removals:

- LanguageModelV1 -> LanguageModel (the v7 type union admits V2/V3/V4 only)
- streamText: maxSteps -> stopWhen(stepCountIs), maxTokens ->
  maxOutputTokens, toolCallStreaming removed (now unconditional)
- usage: promptTokens/completionTokens -> inputTokens/outputTokens
- step.reasoning is an array; use step.reasoningText for the text
- stream parts: textDelta -> text, step-finish -> finish-step
- zod 4 requires an explicit key schema for z.record()

Three breakages that neither typecheck nor the suite would have caught:

- @ai-sdk/openai v2 switched the bare `openai(id)` call to the Responses
  API, which Ollama/vLLM/DeepSeek/minimax do not implement. Pinned to
  `.chat(id)` so every OpenAI-compatible provider stays on
  /chat/completions, with a test asserting the request URL.
- Tools carry their schema in `inputSchema`, not `parameters`. A missing
  inputSchema is not rejected: asSchema(undefined) yields an empty object
  schema, so every tool would have reached the model stripped of its
  arguments. Converted at the single toAiTool boundary; the internal
  `parameters` convention is unchanged.
- Conversation history rebuilt by events-to-messages used v4 part shapes
  (tool-call.args, tool-result.result, image.mimeType). Any second turn
  after a tool call would fail message validation.

reasoning_effort now ships as a call-time provider option, since the
provider dropped the constructor settings argument that used to carry it.
Added tests that assert it reaches the request body, and that it is absent
when unset -- the previous test only checked the config object, so a
silently dropped value would still have passed.

MCP talks to @modelcontextprotocol/sdk (already a dependency) directly,
because the AI SDK removed its experimental_createMCPClient wrapper. The
reconnect/backoff/degradation logic and the mcp_<server>_<tool> namespacing
are untouched.

Behavior change worth noting: an unusable tool call (malformed JSON,
unknown tool) is now reported back to the model instead of throwing, so the
loop recovers and the turn ends idle rather than failed. The call is still
never announced or executed; two status assertions were updated to match.

typecheck (src/tests/console), build, package:check and smoke:release all
pass. Remaining suite failures are pre-existing platform gaps unrelated to
these packages (tar CLI, /bin/sh, POSIX path assertions, symlink EPERM).

TypeScript 7 (PR #100) is intentionally left for a separate commit.

* chore(build): migrate tsup to tsdown, bump typescript to v7
This commit is contained in:
Yao
2026-09-13 12:29:46 +08:00
committed by GitHub
parent a634eb4314
commit 89ae80902f
26 changed files with 3345 additions and 2507 deletions
+3043 -2351
View File
File diff suppressed because it is too large Load Diff
+9 -9
View File
@@ -55,7 +55,7 @@
"dev": "tsx src/index.ts",
"dev:console": "vite --host 0.0.0.0 --config apps/console/vite.config.ts",
"build": "npm run build:runtime && npm run build:console",
"build:runtime": "tsup",
"build:runtime": "tsdown",
"build:console": "npm run typecheck:console && vite build --config apps/console/vite.config.ts",
"prepare": "node scripts/prepare-install.mjs",
"start": "node dist/index.js",
@@ -111,31 +111,31 @@
],
"license": "Apache-2.0",
"dependencies": {
"@ai-sdk/anthropic": "^1.0.0",
"@ai-sdk/openai": "^1.0.0",
"@ai-sdk/anthropic": "^4.0.0",
"@ai-sdk/openai": "^4.0.0",
"@hono/node-server": "^2.1.1",
"@modelcontextprotocol/sdk": "^1.30.0",
"ai": "^4.0.0",
"ai": "^7.0.0",
"commander": "^15.0.0",
"hono": "^4.6.0",
"nanoid": "^6.0.1",
"prefix-safe-json": "0.4.3",
"yaml": "^2.5.0",
"zod": "^3.23.0"
"zod": "^4.0.0"
},
"devDependencies": {
"@types/node": "^26.4.0",
"@types/react": "^19.2.18",
"@types/react-dom": "^19.2.3",
"@vitejs/plugin-react": "^4.7.0",
"@vitejs/plugin-react": "^6.0.0",
"fast-check": "^3.20.0",
"lucide-react": "^1.34.0",
"react": "^19.2.8",
"react-dom": "^19.2.8",
"tsup": "^8.0.0",
"tsdown": "^0.22.14",
"tsx": "^4.23.12",
"typescript": "^5.6.0",
"vite": "^5.4.21",
"typescript": "^7.0.2",
"vite": "^8.0.0",
"vitest": "^2.1.0"
}
}
+2 -2
View File
@@ -24,7 +24,7 @@ const mcpServerConfigSchema = z.discriminatedUnion('type', [
name: z.string().min(1, 'MCP server name is required'),
command: z.string().min(1, 'stdio MCP server command is required'),
args: z.array(z.string()).optional(),
env: z.record(z.string()).optional(),
env: z.record(z.string(), z.string()).optional(),
timeout: z.number().positive().optional(),
}),
]);
@@ -97,7 +97,7 @@ export const agentDefinitionSchema = z.object({
skills: z.array(skillRefSchema).optional(),
mcp_servers: z.array(mcpServerConfigSchema).optional(),
tools: z.array(agentToolsetSchema).optional(),
metadata: z.record(z.unknown()).optional(),
metadata: z.record(z.string(), z.unknown()).optional(),
max_turns: z.number().int().positive().max(1000).optional(),
temperature: z.number().min(0).max(2).optional(),
delegations: z.array(z.string()).optional(),
+29 -15
View File
@@ -2,9 +2,9 @@
* MCP Client Manager (Requirement 5)
*
* Connects to MCP servers declared in an Agent definition and exposes their
* tools to the engine loop. Uses the Vercel AI SDK's built-in MCP client
* (`experimental_createMCPClient`) so the returned tools plug straight into
* `streamText`.
* tools to the engine loop. Uses the official MCP SDK client
* (`@modelcontextprotocol/sdk`), since the Vercel AI SDK removed the
* `experimental_createMCPClient` wrapper this used to go through.
*
* Transports:
* - stdio: spawns a subprocess and speaks MCP over stdin/stdout
@@ -15,8 +15,9 @@
* whatever tools did connect.
*/
import { experimental_createMCPClient } from 'ai';
import { Experimental_StdioMCPTransport } from 'ai/mcp-stdio';
import { Client } from '@modelcontextprotocol/sdk/client/index.js';
import { StdioClientTransport } from '@modelcontextprotocol/sdk/client/stdio.js';
import { SSEClientTransport } from '@modelcontextprotocol/sdk/client/sse.js';
import { resolveEnvVarsDeep } from '@/core/config/env-resolver.js';
import type { McpServerConfig } from '@/types/agent.js';
@@ -212,27 +213,40 @@ export class McpManager {
// ============================================================
private async createClient(server: McpServerConfig): Promise<McpClient> {
let transport;
if (server.type === 'stdio') {
if (!server.command) {
throw new Error(`MCP server "${server.name}": stdio transport requires "command"`);
}
const env = server.env ? resolveEnvVarsDeep(server.env, false) : undefined;
const transport = new Experimental_StdioMCPTransport({
transport = new StdioClientTransport({
command: server.command,
args: server.args ?? [],
env,
});
return experimental_createMCPClient({ transport }) as unknown as Promise<McpClient>;
} else {
if (!server.url) {
throw new Error(`MCP server "${server.name}": url transport requires "url"`);
}
const url = resolveEnvVarsDeep(server.url, false);
transport = new SSEClientTransport(new URL(url));
}
// url transport
if (!server.url) {
throw new Error(`MCP server "${server.name}": url transport requires "url"`);
}
const url = resolveEnvVarsDeep(server.url, false);
return experimental_createMCPClient({
transport: { type: 'sse', url },
}) as unknown as Promise<McpClient>;
const client = new Client({ name: 'sandbase-harness', version: '1.0.0' });
await client.connect(transport);
return {
async tools() {
const { tools } = await client.listTools();
return Object.fromEntries(tools.map((tool) => [tool.name, {
description: tool.description,
parameters: tool.inputSchema,
execute: (args: unknown) =>
client.callTool({ name: tool.name, arguments: (args ?? {}) as Record<string, unknown> }),
}]));
},
close: () => client.close(),
};
}
}
+2 -2
View File
@@ -11,7 +11,7 @@
* decide when to trigger without pulling in a tokenizer dependency.
*/
import { generateText, type LanguageModelV1 } from 'ai';
import { generateText, type LanguageModel } from 'ai';
import type { Message } from './events-to-messages.js';
/** Fire compaction when estimated tokens exceed this fraction of the window. */
@@ -61,7 +61,7 @@ export class ContextCompactor {
* summary string via the model. Returns null if there's nothing worth
* compacting (too few messages).
*/
async compact(messages: Message[], model: LanguageModelV1): Promise<CompactionResult | null> {
async compact(messages: Message[], model: LanguageModel): Promise<CompactionResult | null> {
const preserveTail = this.config.preserveTailMessages ?? PRESERVE_TAIL_MESSAGES;
if (messages.length <= preserveTail + 1) {
return null; // not enough history to bother
+10 -6
View File
@@ -44,18 +44,20 @@ export interface SystemMessage {
export type ContentPart =
| { type: 'text'; text: string }
| { type: 'image'; image: string; mimeType?: string };
| { type: 'image'; image: string; mediaType?: string };
export type AssistantContentPart =
| { type: 'text'; text: string }
| { type: 'reasoning'; text: string; providerOptions?: Record<string, unknown> }
| { type: 'tool-call'; toolCallId: string; toolName: string; args: Record<string, unknown> };
| { type: 'tool-call'; toolCallId: string; toolName: string; input: Record<string, unknown> };
export type ToolResultPart = {
type: 'tool-result';
toolCallId: string;
toolName: string;
result: unknown;
output:
| { type: 'text'; value: string }
| { type: 'json'; value: unknown };
};
export type Message = UserMessage | AssistantMessage | ToolMessage | SystemMessage;
@@ -179,7 +181,7 @@ export function eventsToMessages(
type: 'tool-call',
toolCallId: block.id,
toolName: block.name,
args: block.input,
input: block.input,
});
}
break;
@@ -198,7 +200,9 @@ export function eventsToMessages(
type: 'tool-result',
toolCallId: resultBlock.tool_use_id,
toolName,
result: resultBlock.content,
output: typeof resultBlock.content === 'string'
? { type: 'text', value: resultBlock.content }
: { type: 'json', value: resultBlock.content },
});
}
break;
@@ -249,7 +253,7 @@ function userContentToParts(content?: ContentBlock[]): ContentPart[] {
} else if (block.type === 'image' && 'source' in block) {
const src = (block as any).source;
if (src?.data) {
parts.push({ type: 'image', image: src.data, mimeType: src.media_type });
parts.push({ type: 'image', image: src.data, mediaType: src.media_type });
} else if (src?.url) {
parts.push({ type: 'text', text: `[Image: ${src.url}]` });
}
+1 -1
View File
@@ -1,7 +1,7 @@
import { z } from 'zod';
import { MINIMAX_PROVIDER } from '@/core/model/minimax.js';
const optionsSchema = z.record(z.unknown()).default({});
const optionsSchema = z.record(z.string(), z.unknown()).default({});
const STORED_SECRET_PREFIX = '__managed_secret__:';
export const runtimeSettingsSchema = z.object({
+33 -13
View File
@@ -8,7 +8,7 @@
import { createOpenAI } from '@ai-sdk/openai';
import { createAnthropic } from '@ai-sdk/anthropic';
import { wrapLanguageModel, type LanguageModelV1, type LanguageModelV1Middleware } from 'ai';
import { wrapLanguageModel, type LanguageModel, type LanguageModelMiddleware } from 'ai';
import { resolveEnvVars } from '@/core/config/env-resolver.js';
import { MINIMAX_PROVIDER, miniMaxOpenAiBaseUrl } from '@/core/model/minimax.js';
import {
@@ -121,7 +121,7 @@ export class ModelRegistry {
* Create a Vercel AI SDK LanguageModel instance, wrapped with the retry
* middleware (Property 14). Resolves ${ENV_VAR} in api_key and base_url.
*/
createModel(name: string): LanguageModelV1 {
createModel(name: string): LanguageModel {
const config = this.resolveModelConfig(name);
if (!config.model) {
throw new ModelNotFoundError(name, Array.from(this.models.keys()), 'Agent model id is required.');
@@ -134,12 +134,13 @@ export class ModelRegistry {
config.model,
resolvedApiKey,
resolvedBaseUrl,
config.reasoning_effort,
);
return wrapLanguageModel({
model: base,
middleware: createRetryMiddleware(this.retryPolicy),
});
const middleware: LanguageModelMiddleware[] = [createRetryMiddleware(this.retryPolicy)];
// Only the OpenAI-compatible branches ever took a reasoning effort.
if (config.reasoning_effort && config.provider !== 'anthropic' && config.provider !== MINIMAX_PROVIDER) {
middleware.push(createReasoningEffortMiddleware(config.reasoning_effort));
}
return wrapLanguageModel({ model: base, middleware });
}
/**
@@ -237,7 +238,7 @@ function publicBaseUrl(value?: string): string | undefined {
* - rate limit (429): honor Retry-After, up to 3x
* - auth (401/403): never retry
*/
function createRetryMiddleware(policy: RetryPolicy): LanguageModelV1Middleware {
function createRetryMiddleware(policy: RetryPolicy): LanguageModelMiddleware {
const runWithRetry = async <T>(fn: () => PromiseLike<T>): Promise<T> => {
let attempt = 0;
for (;;) {
@@ -263,6 +264,22 @@ function createRetryMiddleware(policy: RetryPolicy): LanguageModelV1Middleware {
};
}
/**
* Pass `reasoning_effort` as a call-time provider option: the provider dropped
* the constructor settings argument that used to carry it.
*/
function createReasoningEffortMiddleware(reasoningEffort: string): LanguageModelMiddleware {
return {
transformParams: async ({ params }) => ({
...params,
providerOptions: {
...params.providerOptions,
openai: { reasoningEffort, ...params.providerOptions?.openai },
},
}),
};
}
function extractHeaders(err: unknown): Headers | undefined {
if (err && typeof err === 'object' && 'responseHeaders' in err) {
const h = (err as { responseHeaders?: unknown }).responseHeaders;
@@ -282,13 +299,16 @@ function sleep(ms: number): Promise<void> {
// Model Factory
// ============================================================
/**
* `.chat(model)` not `openai(model)`: the bare call now targets the Responses
* API, which Ollama/vLLM/DeepSeek/minimax do not implement.
*/
function createModelInstance(
provider: ModelProviderType,
model: string,
apiKey?: string,
baseUrl?: string,
reasoningEffort?: string,
): LanguageModelV1 {
) {
switch (provider) {
case 'openai':
case 'ollama': {
@@ -296,14 +316,14 @@ function createModelInstance(
apiKey: apiKey ?? 'ollama', // Ollama doesn't need a key
baseURL: baseUrl,
});
return openai(model, reasoningEffort ? { reasoningEffort: reasoningEffort as 'low' | 'medium' | 'high' } : undefined);
return openai.chat(model);
}
case MINIMAX_PROVIDER: {
const minimax = createOpenAI({
apiKey: apiKey ?? '',
baseURL: miniMaxOpenAiBaseUrl({}, baseUrl),
});
return minimax(model);
return minimax.chat(model);
}
case 'anthropic': {
const anthropic = createAnthropic({
@@ -318,7 +338,7 @@ function createModelInstance(
apiKey: apiKey ?? '',
baseURL: baseUrl,
});
return openaiCompat(model, reasoningEffort ? { reasoningEffort: reasoningEffort as 'low' | 'medium' | 'high' } : undefined);
return openaiCompat.chat(model);
}
}
}
+1
View File
@@ -1,3 +1,4 @@
/// <reference types="node" />
/**
* managed-agents SDK — public client entry point.
*
+23 -18
View File
@@ -14,9 +14,9 @@
* Reference: OMA default-loop.ts
*/
import { jsonSchema, streamText } from 'ai';
import { jsonSchema, stepCountIs, streamText } from 'ai';
import { createAiSdkExecutionLock, type JsonSchemaLike } from 'prefix-safe-json';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
import type { AgentStrategy, StrategyContext } from '@/types/strategy.js';
import type { SessionEvent } from '@/types/session.js';
import type { ContentBlock } from '@/types/cma-protocol.js';
@@ -168,20 +168,19 @@ export class DefaultStrategy implements AgentStrategy {
}));
const result = streamText({
model: model as LanguageModelV1,
model: model as LanguageModel,
system: systemPrompt || undefined,
messages: aiMessages,
tools: Object.keys(aiTools).length > 0 ? aiTools : undefined,
maxSteps,
stopWhen: stepCountIs(maxSteps),
temperature: config.temperature,
maxTokens: config.maxTokens,
toolCallStreaming: true,
maxOutputTokens: config.maxTokens,
abortSignal,
onStepFinish: async (step) => {
totalSteps++;
const tokensIn = step.usage?.promptTokens ?? 0;
const tokensOut = step.usage?.completionTokens ?? 0;
const tokensIn = step.usage?.inputTokens ?? 0;
const tokensOut = step.usage?.outputTokens ?? 0;
totalTokensIn += tokensIn;
totalTokensOut += tokensOut;
@@ -199,7 +198,7 @@ export class DefaultStrategy implements AgentStrategy {
eventLog.recordUsage(session.id, tokensIn, tokensOut);
// Emit agent.thinking for reasoning output (extended-thinking models)
const reasoning = (step as { reasoning?: string }).reasoning;
const reasoning = step.reasoningText;
if (reasoning && reasoning.trim()) {
const thinkingEvent = eventLog.append(session.id, {
type: 'agent.thinking',
@@ -242,7 +241,7 @@ export class DefaultStrategy implements AgentStrategy {
type: 'tool_use',
id: toolCall.toolCallId,
name: toolCall.toolName,
input: toolCall.args as Record<string, unknown>,
input: toolCall.input as Record<string, unknown>,
}] as ContentBlock[],
tokensIn,
tokensOut,
@@ -255,9 +254,9 @@ export class DefaultStrategy implements AgentStrategy {
if (step.toolResults && step.toolResults.length > 0) {
for (const toolResult of step.toolResults) {
const isMcp = toolResult.toolName?.startsWith('mcp_') ?? false;
const raw = typeof toolResult.result === 'string'
? toolResult.result
: JSON.stringify(toolResult.result);
const raw = typeof toolResult.output === 'string'
? toolResult.output
: JSON.stringify(toolResult.output);
const capped = raw.length > MAX_TOOL_RESULT_CHARS
? raw.slice(0, MAX_TOOL_RESULT_CHARS) + `\n\n[truncated: ${raw.length - MAX_TOOL_RESULT_CHARS} more chars]`
: raw;
@@ -305,10 +304,10 @@ export class DefaultStrategy implements AgentStrategy {
broadcast(
transientEvent(session.id, 'agent.message_chunk', {
message_id: messageId,
delta: part.textDelta,
delta: part.text,
}),
);
} else if (part.type === 'step-finish' || part.type === 'finish') {
} else if (part.type === 'finish-step' || part.type === 'finish') {
if (streaming) {
broadcast(transientEvent(session.id, 'agent.message_stream_end', { message_id: messageId }));
streaming = false;
@@ -393,14 +392,20 @@ export class DefaultStrategy implements AgentStrategy {
}
}
/**
* Translate this codebase's `parameters` convention to the SDK's `inputSchema`.
* A tool left on `parameters` is not rejected — it reaches the model with an
* empty argument schema — so this conversion cannot be skipped.
*/
function toAiTool(tool: any): any {
if (!tool || typeof tool !== 'object') return tool;
if (tool.inputSchema) return tool;
if (!tool.parameters || typeof tool.parameters !== 'object') return tool;
if (isAiSdkSchema(tool.parameters)) return tool;
const { parameters, ...rest } = tool;
return {
...tool,
parameters: jsonSchema(tool.parameters),
...rest,
inputSchema: isAiSdkSchema(parameters) ? parameters : jsonSchema(parameters),
};
}
+3 -3
View File
@@ -2,10 +2,10 @@
* Model Provider Types
*
* Unified interface for model providers (OpenAI, Anthropic, Ollama, vLLM, etc.)
* All providers are abstracted through the Vercel AI SDK LanguageModelV1 interface.
* All providers are abstracted through the Vercel AI SDK LanguageModel interface.
*/
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
// ============================================================
// Model Provider Interface
@@ -16,7 +16,7 @@ export interface ModelProvider {
readonly type: ModelProviderType;
/** Create a Vercel AI SDK-compatible LanguageModel instance */
createModel(config: ModelConfig): LanguageModelV1;
createModel(config: ModelConfig): LanguageModel;
/** Health check — returns false on any error, does not throw */
healthCheck(): Promise<boolean>;
+1
View File
@@ -1,3 +1,4 @@
/// <reference types="node" />
/**
* Sandbox Provider Types
*
+2 -2
View File
@@ -5,7 +5,7 @@
* context building → LLM call → response parsing → tool execution → loop.
*/
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
import type { SandboxInstance } from './sandbox.js';
import type { Session, SessionEvent } from './session.js';
@@ -33,7 +33,7 @@ export interface StrategyContext {
/** Agent system prompt (with any injected skills). Sent to the model. */
systemPrompt: string;
messages: CoreMessage[];
model: LanguageModelV1;
model: LanguageModel;
tools: Record<string, CoreTool>;
sandbox: SandboxInstance;
eventLog: EventLogWriter;
+9 -8
View File
@@ -15,7 +15,7 @@ import { ModelRegistry } from '@/model/registry.js';
import { LocalSandboxProvider } from '@/sandbox/local-provider.js';
import { eventsToMessages } from '@/core/session/events-to-messages.js';
import type { AgentStrategy, StrategyContext } from '@/types/strategy.js';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
// Strategy that just records the message count it was handed and emits nothing.
class NoopStrategy implements AgentStrategy {
@@ -28,21 +28,22 @@ class NoopStrategy implements AgentStrategy {
}
}
function fakeModel(): LanguageModelV1 {
function fakeModel(): LanguageModel {
return {
specificationVersion: 'v1',
specificationVersion: 'v4',
provider: 'test',
modelId: 'test',
supportedUrls: {},
async doGenerate() {
return {
text: 'SUMMARY: prior conversation compacted',
finishReason: 'stop',
usage: { promptTokens: 1, completionTokens: 1 },
rawCall: { rawPrompt: null, rawSettings: {} },
content: [{ type: 'text', text: 'SUMMARY: prior conversation compacted' }],
finishReason: { unified: 'stop', raw: 'stop' },
usage: { inputTokens: { total: 1 }, outputTokens: { total: 1 } },
warnings: [],
} as any;
},
async doStream() { throw new Error('unused'); },
} as unknown as LanguageModelV1;
} as unknown as LanguageModel;
}
describe('Compaction during execution', () => {
+6 -5
View File
@@ -16,7 +16,7 @@ import { LocalSandboxProvider } from '@/sandbox/local-provider.js';
import { InMemoryEventLog } from '@/core/session/in-memory-event-log.js';
import type { AgentStrategy, StrategyContext } from '@/types/strategy.js';
import type { AgentDefinition } from '@/types/agent.js';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
/** Strategy that emits an agent.message echoing the last user text. */
class EchoStrategy implements AgentStrategy {
@@ -37,12 +37,13 @@ class EchoStrategy implements AgentStrategy {
}
}
function fakeModel(): LanguageModelV1 {
function fakeModel(): LanguageModel {
return {
specificationVersion: 'v1', provider: 'test', modelId: 't',
async doGenerate() { return { text: '', finishReason: 'stop', usage: {}, rawCall: { rawPrompt: null, rawSettings: {} } } as any; },
specificationVersion: 'v4', provider: 'test', modelId: 't',
supportedUrls: {},
async doGenerate() { return { content: [], finishReason: { unified: 'stop', raw: 'stop' }, usage: {}, warnings: [] } as any; },
async doStream() { throw new Error('unused'); },
} as unknown as LanguageModelV1;
} as unknown as LanguageModel;
}
describe('Multi-agent delegation', () => {
+6 -5
View File
@@ -18,14 +18,15 @@ import { SqliteMemoryProvider } from '@/core/memory/sqlite-memory-provider.js';
import { ModelRegistry } from '@/model/registry.js';
import { LocalSandboxProvider } from '@/sandbox/local-provider.js';
import type { AgentStrategy, StrategyContext } from '@/types/strategy.js';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
function fakeModel(): LanguageModelV1 {
function fakeModel(): LanguageModel {
return {
specificationVersion: 'v1', provider: 'test', modelId: 't',
async doGenerate() { return { text: '', finishReason: 'stop', usage: {}, rawCall: { rawPrompt: null, rawSettings: {} } } as any; },
specificationVersion: 'v4', provider: 'test', modelId: 't',
supportedUrls: {},
async doGenerate() { return { content: [], finishReason: { unified: 'stop', raw: 'stop' }, usage: {}, warnings: [] } as any; },
async doStream() { throw new Error('unused'); },
} as unknown as LanguageModelV1;
} as unknown as LanguageModel;
}
class CapturingStrategy implements AgentStrategy {
+6 -5
View File
@@ -14,14 +14,15 @@ import { SnapshotManager } from '@/core/session/snapshot-manager.js';
import { ModelRegistry } from '@/model/registry.js';
import { LocalSandboxProvider } from '@/sandbox/local-provider.js';
import type { AgentStrategy, StrategyContext } from '@/types/strategy.js';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
function fakeModel(): LanguageModelV1 {
function fakeModel(): LanguageModel {
return {
specificationVersion: 'v1', provider: 'test', modelId: 't',
async doGenerate() { return { text: '', finishReason: 'stop', usage: {}, rawCall: { rawPrompt: null, rawSettings: {} } } as any; },
specificationVersion: 'v4', provider: 'test', modelId: 't',
supportedUrls: {},
async doGenerate() { return { content: [], finishReason: { unified: 'stop', raw: 'stop' }, usage: {}, warnings: [] } as any; },
async doStream() { throw new Error('unused'); },
} as unknown as LanguageModelV1;
} as unknown as LanguageModel;
}
/** Strategy that writes a file into the sandbox, then (on a later turn) reads it. */
+5 -4
View File
@@ -13,21 +13,22 @@ import { DefaultSessionExecutor } from '@/core/session/executor.js';
import { DefaultStrategy } from '@/strategy/default-strategy.js';
import { ModelRegistry } from '@/model/registry.js';
import { LocalSandboxProvider } from '@/sandbox/local-provider.js';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
/** A model whose stream throws — simulates a provider/auth/network failure. */
function throwingModel(): LanguageModelV1 {
function throwingModel(): LanguageModel {
return {
specificationVersion: 'v1',
specificationVersion: 'v4',
provider: 'test',
modelId: 'boom',
supportedUrls: {},
async doGenerate() {
throw new Error('401 unauthorized');
},
async doStream() {
throw new Error('401 unauthorized');
},
} as unknown as LanguageModelV1;
} as unknown as LanguageModel;
}
describe('Model error → failed turn', () => {
+16 -15
View File
@@ -8,38 +8,39 @@ import { DefaultSessionExecutor } from '@/core/session/executor.js';
import { DefaultStrategy } from '@/strategy/default-strategy.js';
import { ModelRegistry } from '@/model/registry.js';
import { LocalSandboxProvider } from '@/sandbox/local-provider.js';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
function streamingTextModel(): LanguageModelV1 {
const USAGE = { inputTokens: { total: 1 }, outputTokens: { total: 1 } };
const STOP = { unified: 'stop', raw: 'stop' } as const;
function streamingTextModel(): LanguageModel {
return {
specificationVersion: 'v1',
specificationVersion: 'v4',
provider: 'test',
modelId: 'streaming-text',
supportedUrls: {},
async doGenerate() {
return {
text: 'ok',
finishReason: 'stop',
usage: { promptTokens: 1, completionTokens: 1 },
rawCall: { rawPrompt: null, rawSettings: {} },
content: [{ type: 'text', text: 'ok' }],
finishReason: STOP,
usage: USAGE,
warnings: [],
} as any;
},
async doStream() {
return {
stream: new ReadableStream({
start(controller) {
controller.enqueue({ type: 'text-delta', textDelta: 'ok' });
controller.enqueue({
type: 'finish',
finishReason: 'stop',
usage: { promptTokens: 1, completionTokens: 1 },
});
controller.enqueue({ type: 'text-start', id: 'txt_1' });
controller.enqueue({ type: 'text-delta', id: 'txt_1', delta: 'ok' });
controller.enqueue({ type: 'text-end', id: 'txt_1' });
controller.enqueue({ type: 'finish', finishReason: STOP, usage: USAGE });
controller.close();
},
}),
rawCall: { rawPrompt: null, rawSettings: {} },
} as any;
},
} as unknown as LanguageModelV1;
} as unknown as LanguageModel;
}
describe('DefaultStrategy tool schemas', () => {
@@ -2,7 +2,7 @@ import { afterEach, describe, expect, it } from 'vitest';
import { mkdtempSync, rmSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
import { Database } from '@/core/db/database.js';
import { ModelRegistry } from '@/model/registry.js';
import { LocalSandboxProvider } from '@/sandbox/local-provider.js';
@@ -20,13 +20,16 @@ type ToolStream = {
rawLifecycle?: boolean;
};
function scriptedToolModel(script: ToolStream): LanguageModelV1 {
const USAGE = { inputTokens: { total: 1 }, outputTokens: { total: 1 } };
function scriptedToolModel(script: ToolStream): LanguageModel {
let turn = 0;
return {
specificationVersion: 'v1',
specificationVersion: 'v4',
provider: 'test',
modelId: 'scripted-tool-call',
supportedUrls: {},
async doGenerate() {
throw new Error('not used');
},
@@ -36,42 +39,44 @@ function scriptedToolModel(script: ToolStream): LanguageModelV1 {
stream: new ReadableStream({
start(controller) {
if (turn === 1) {
const toolName = script.toolName ?? 'write';
// Raw argument bytes now arrive as the tool-input lifecycle.
if (script.rawLifecycle !== false) {
controller.enqueue({ type: 'tool-input-start', id: 'call_write_1', toolName });
controller.enqueue({
type: 'tool-call-delta',
toolCallType: 'function',
toolCallId: 'call_write_1',
toolName: script.toolName ?? 'write',
argsTextDelta: script.rawArgs ?? script.args,
type: 'tool-input-delta',
id: 'call_write_1',
delta: script.rawArgs ?? script.args,
});
controller.enqueue({ type: 'tool-input-end', id: 'call_write_1' });
}
controller.enqueue({
type: 'tool-call',
toolCallType: 'function',
toolCallId: 'call_write_1',
toolName: script.toolName ?? 'write',
args: script.args,
toolName,
input: script.args,
});
controller.enqueue({
type: 'finish',
finishReason: script.finishReason,
usage: { promptTokens: 1, completionTokens: 1 },
finishReason: { unified: script.finishReason, raw: script.finishReason },
usage: USAGE,
});
} else {
controller.enqueue({ type: 'text-delta', textDelta: 'continued' });
controller.enqueue({ type: 'text-start', id: 'txt_1' });
controller.enqueue({ type: 'text-delta', id: 'txt_1', delta: 'continued' });
controller.enqueue({ type: 'text-end', id: 'txt_1' });
controller.enqueue({
type: 'finish',
finishReason: 'stop',
usage: { promptTokens: 1, completionTokens: 1 },
finishReason: { unified: 'stop', raw: 'stop' },
usage: USAGE,
});
}
controller.close();
},
}),
rawCall: { rawPrompt: null, rawSettings: {} },
} as any;
},
} as unknown as LanguageModelV1;
} as unknown as LanguageModel;
}
class CountingLocalSandboxProvider extends LocalSandboxProvider {
@@ -213,10 +218,13 @@ describe('tool confirmation stream reliability', () => {
expect(harness.sandboxProvider.writeCount).toBe(0);
});
// An unusable tool call is now reported back to the model instead of throwing,
// so the turn ends idle rather than failed (same as the schema-invalid case
// below). What must not happen is the call being announced or executed.
it('does not execute malformed JSON', async () => {
const harness = createHarness({ finishReason: 'tool-calls', args: '{"path":' });
const sessionId = await requestToolCall(harness);
await waitFor(() => harness.manager.get(sessionId)?.status === 'failed');
await waitFor(() => harness.manager.get(sessionId)?.status === 'paused');
expect(harness.manager.getEventLogger().getEvents(sessionId)
.filter((event) => event.type === 'agent.tool_use')).toHaveLength(0);
@@ -243,7 +251,7 @@ describe('tool confirmation stream reliability', () => {
args: '{}',
});
const sessionId = await requestToolCall(harness);
await waitFor(() => harness.manager.get(sessionId)?.status === 'failed');
await waitFor(() => harness.manager.get(sessionId)?.status === 'paused');
expect(harness.sandboxProvider.writeCount).toBe(0);
});
+6 -5
View File
@@ -17,14 +17,15 @@ import { DefaultSessionExecutor } from '@/core/session/executor.js';
import { ModelRegistry } from '@/model/registry.js';
import { LocalSandboxProvider } from '@/sandbox/local-provider.js';
import type { AgentStrategy, StrategyContext } from '@/types/strategy.js';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
function fakeModel(): LanguageModelV1 {
function fakeModel(): LanguageModel {
return {
specificationVersion: 'v1', provider: 'test', modelId: 't',
async doGenerate() { return { text: '', finishReason: 'stop', usage: {}, rawCall: { rawPrompt: null, rawSettings: {} } } as any; },
specificationVersion: 'v4', provider: 'test', modelId: 't',
supportedUrls: {},
async doGenerate() { return { content: [], finishReason: { unified: 'stop', raw: 'stop' }, usage: {}, warnings: [] } as any; },
async doStream() { throw new Error('unused'); },
} as unknown as LanguageModelV1;
} as unknown as LanguageModel;
}
describe('Tool confirmation — requires_action transition', () => {
+9 -9
View File
@@ -8,31 +8,31 @@ import {
estimateMessagesTokens,
} from '@/core/session/context-compactor.js';
import type { Message } from '@/core/session/events-to-messages.js';
import type { LanguageModelV1 } from 'ai';
import type { LanguageModel } from 'ai';
function userMsg(text: string): Message {
return { role: 'user', content: [{ type: 'text', text }] };
}
/** A fake model that returns a fixed summary via generateText. */
function fakeModel(summaryText: string): LanguageModelV1 {
function fakeModel(summaryText: string): LanguageModel {
return {
specificationVersion: 'v1',
specificationVersion: 'v4',
provider: 'test',
modelId: 'test-model',
defaultObjectGenerationMode: undefined,
supportedUrls: {},
async doGenerate() {
return {
text: summaryText,
finishReason: 'stop',
usage: { promptTokens: 10, completionTokens: 5 },
rawCall: { rawPrompt: null, rawSettings: {} },
content: [{ type: 'text', text: summaryText }],
finishReason: { unified: 'stop', raw: 'stop' },
usage: { inputTokens: { total: 10 }, outputTokens: { total: 5 } },
warnings: [],
} as any;
},
async doStream() {
throw new Error('not used');
},
} as unknown as LanguageModelV1;
} as unknown as LanguageModel;
}
describe('ContextCompactor', () => {
+82 -1
View File
@@ -1,4 +1,4 @@
import { describe, expect, it } from 'vitest';
import { afterEach, describe, expect, it, vi } from 'vitest';
import { ModelRegistry } from '@/model/registry.js';
describe('ModelRegistry runtime introspection', () => {
@@ -129,3 +129,84 @@ describe('ModelRegistry runtime introspection', () => {
});
});
});
// Guards the wiring, not just the config object: asserting only that
// resolveModelConfig keeps the field would pass even if it never reached the model.
describe('ModelRegistry reasoning effort wiring', () => {
afterEach(() => {
vi.unstubAllGlobals();
});
/** Capture the request the provider would send. */
function stubFetch(): { body: () => Record<string, unknown>; url: () => string } {
let captured: Record<string, unknown> = {};
let url = '';
vi.stubGlobal('fetch', async (requestUrl: unknown, init: { body?: string }) => {
url = String(requestUrl);
captured = JSON.parse(init?.body ?? '{}');
return new Response(
JSON.stringify({
id: 'x',
created: 0,
model: 'm',
choices: [{ index: 0, message: { role: 'assistant', content: 'ok' }, finish_reason: 'stop' }],
usage: { prompt_tokens: 1, completion_tokens: 1 },
}),
{ status: 200, headers: { 'content-type': 'application/json' } },
);
});
return { body: () => captured, url: () => url };
}
const prompt = [{ role: 'user' as const, content: [{ type: 'text' as const, text: 'hi' }] }];
it('sends the configured reasoning effort to the provider', async () => {
const registry = new ModelRegistry();
registry.register({
name: 'default',
provider: 'openai_compatible',
api_key: 'test-key',
base_url: 'https://example.invalid/v1',
reasoning_effort: 'high',
is_default: true,
});
const fetchStub = stubFetch();
await (registry.createModel('deepseek-v4-pro') as any).doGenerate({ prompt });
expect(fetchStub.body()).toMatchObject({ reasoning_effort: 'high' });
});
it('omits reasoning effort when none is configured', async () => {
const registry = new ModelRegistry();
registry.register({
name: 'default',
provider: 'openai_compatible',
api_key: 'test-key',
base_url: 'https://example.invalid/v1',
is_default: true,
});
const fetchStub = stubFetch();
await (registry.createModel('deepseek-v4-pro') as any).doGenerate({ prompt });
expect(fetchStub.body()).not.toHaveProperty('reasoning_effort');
});
// DeepSeek/Ollama/vLLM/minimax implement /chat/completions but not /responses.
it('targets the chat completions endpoint for OpenAI-compatible providers', async () => {
const registry = new ModelRegistry();
registry.register({
name: 'default',
provider: 'openai_compatible',
api_key: 'test-key',
base_url: 'https://example.invalid/v1',
is_default: true,
});
const fetchStub = stubFetch();
await (registry.createModel('deepseek-v4-pro') as any).doGenerate({ prompt });
expect(fetchStub.url()).toBe('https://example.invalid/v1/chat/completions');
});
});
+2 -1
View File
@@ -413,9 +413,10 @@ describe('Settings V2 activation', () => {
const activated = activateRuntimeSettings(db);
expect(activated.activation_status).toBe('failed');
// `code` is relayed verbatim from zod, which now reports this as invalid_value.
expect(activated.activation_errors).toContainEqual(expect.objectContaining({
path: 'schema_version',
code: 'invalid_literal',
code: 'invalid_value',
}));
expect(activated.effective_revision).toBe(1);
expect(activated.effective_config.loop_engine.options.default_max_steps).toBe(25);
+2 -2
View File
@@ -4,6 +4,7 @@
"module": "ESNext",
"moduleResolution": "bundler",
"lib": ["ES2022"],
"types": ["node"],
"outDir": "./dist",
"rootDir": "./src",
"strict": true,
@@ -16,8 +17,7 @@
"sourceMap": true,
"paths": {
"@/*": ["./src/*"]
},
"baseUrl": "."
}
},
"include": ["src/**/*"],
"exclude": ["node_modules", "dist", "tests"]
+9 -5
View File
@@ -1,8 +1,8 @@
import { defineConfig } from 'tsup';
import { defineConfig } from 'tsdown';
import { writeFileSync, readFileSync } from 'node:fs';
export default defineConfig([
{
// CLI entry (needs the shebang)
entry: { index: 'src/index.ts', 'mcp/index': 'src/mcp/index.ts' },
format: ['esm'],
dts: true,
@@ -10,11 +10,10 @@ export default defineConfig([
clean: true,
platform: 'node',
target: 'node22',
splitting: false,
outExtensions: () => ({ js: '.js', dts: '.d.ts' }),
banner: { js: '#!/usr/bin/env node' },
},
{
// SDK entry (library import, no shebang)
entry: { sdk: 'src/sdk/index.ts' },
format: ['esm'],
dts: true,
@@ -22,6 +21,11 @@ export default defineConfig([
clean: false,
platform: 'node',
target: 'node22',
splitting: false,
outExtensions: () => ({ js: '.js', dts: '.d.ts' }),
onSuccess: async () => {
const dtsPath = 'dist/sdk.d.ts';
const content = readFileSync(dtsPath, 'utf-8');
writeFileSync(dtsPath, `/// <reference types="node" />\n${content}`);
},
},
]);