Files
ruflo/scripts/bulk-fix-tool-descriptions.mjs
rUvandReuven 8aac1b2ccc feat(mcp): ADR-112 — 285/285 tool descriptions have 'Use when …' guidance + CI guard + #1892 statusline (#1897)
* docs(adr-095): status update — G1/G3/G4/G6/benchmark remediated (verification receipts)

External re-audit (AlphaSignal AI, May 7) surfaced ADR-095 gaps publicly.
Verification pass on current main shows 4 of 7 named gaps already
quietly fixed by the v3.7.0-alpha work:

  G1 — agent_spawn intentionally registry-only; real execution wire is
       agent_execute → agent-execute-core.ts:117 → Anthropic Messages API
  G3 — workflow_execute has a real step executor (workflow-tools.ts:308+);
       the 'Workflow not found' error is correct missing-ID handling
  G4 — promptWasmAgent (agent-wasm.ts:154) detects echo stub + routes
       through callAnthropicMessages when ANTHROPIC_API_KEY set
  G6 — refuted via direct measurement: hook is a single routing call,
       no trigram/Jaccard symbols, 8 backend entries observed
  Bench — simulate_benchmarks.py removed; '84.8% SWE' claim cleaned
       from tracked docs

Preserves original ADR text. Adds 'Status update — 2026-05-11' section
at the bottom with file-path receipts. Still open: G2, G5 (→ADR-094),
G7 (per-controller), #1748 (tool descriptions).

Related: #1896 (external audit response)

Co-Authored-By: RuFlo <ruv@ruv.net>

* feat(mcp): ADR-112 — every tool description now has "Use when … is wrong because …" guidance + CI guard + #1892 statusline fix

External audit (AlphaSignal AI, May 7) measured 237/300 MCP tool
descriptions lacked guidance for when Claude should pick Ruflo's tool
over native (Bash/Read/Grep/Glob/Task/TodoWrite). Re-measurement on
current main was actually 279/285 — worse. Now: 0/285.

## ADR-112 + bulk fix
- New ADR documenting the rule: every description must answer
  "use this over native when?" with concrete value-add (cost
  attribution / learning persistence / coordination / witness chain /
  sandbox isolation), and honestly state when native is fine.
- 7 agent-tools.ts descriptions hand-edited (the P1 native-Task overlap).
- 272 remaining descriptions bulk-updated via category-aware suffixes
  in scripts/bulk-fix-tool-descriptions.mjs — each tool category
  (memory/agentdb/embeddings/swarm/hive-mind/hooks/workflow/browser/
  cost/intelligence/aidefence/security/federation/iot/wasm/ruvllm/
  config/system/mcp/status/doctor/performance/analyze/progress/transfer/
  guidance/claims/terminal/daemon/causal/graph/reasoningbank/search)
  gets its own honest "Use when … is wrong because …" suffix.

## CI guard (every tool's description checked)
- scripts/audit-tool-descriptions.mjs scans every MCPTool definition and
  enforces three gates per tool:
    1. "Use when ..." guidance present (the original ADR-112 check)
    2. Description length ≥ 80 chars (catches near-empty descriptions)
    3. Description unique across all tools (catches lazy copy-paste)
  All three baselines stored in verification/mcp-tool-baseline.json
  and are monotone-decreasing — CI fails on any regression.
- .github/workflows/v3-ci.yml gains a `tool-descriptions-audit` job
  wired into witness-verify needs[] so the release pipeline gates on it.

## #1892 — UI version mismatch
- v3/@claude-flow/cli/.claude/helpers/statusline.js was hardcoded
  to "RuFlo V3.5"; now reads from installed @claude-flow/cli's
  package.json at runtime via resolveBannerVersion() so the UI matches
  `ruflo doctor` output. Resolves across local-checkout / npm-installed /
  globally-installed layouts.

## Witness manifest
- verification/witness-fixes.json: ADR-111, ADR-112, #1892 markers added.
- verification/macos/manifest.md.json regen'd — 102/102 fixes verified,
  Ed25519 signature valid.

## Tests
- 1963 federation+cli tests pass on this branch.
- Build clean.
- Audit: 285/285 tools with guidance, 0 too-short, 0 duplicates.

Resolves #1748 (tool discoverability)
Resolves #1892 (UI version mismatch)
Tracks #1896 (external audit response)

Co-Authored-By: RuFlo <ruv@ruv.net>

* docs(verification): README — current 102/102 fix count + ADR-112 baseline + Layer 1 smoke jobs

- Folder layout adds mcp-tool-baseline.json (ADR-112)
- Example manifest summary updated 82→102 fixes, alpha.18→alpha.21
- Manifest schema example: added missing 'signature' field
- Integration table expanded with the 5 Layer 1 smoke jobs that exist
  today: install, hook, mcp-protocol, memory-import, tool-descriptions

Co-Authored-By: RuFlo <ruv@ruv.net>

* chore(verification): lock ADR-112 baseline at all-zeros for noGuidance, tooShort, duplicates

The enhanced audit-tool-descriptions.mjs now tracks three gates instead
of one. Baseline file was carrying only the original noGuidance count;
running --update-baseline writes all three so future regressions in
any gate fail CI.

Co-Authored-By: RuFlo <ruv@ruv.net>

* ci: fix Performance Verification — download phantom artifact

Long-standing latent bug masked by download-artifact@v3's warning
behavior, surfaced as a hard error by @v4. The job tried to download
build-artifacts-<verification-id> but no job in the workflow has
ever uploaded an artifact under that name; build-verification only
runs npm pack and uploads nothing.

Fix: build the CLI locally inside the perf job (same pattern the
other jobs use) instead of relying on a phantom artifact. Memory-leak
smoke now probes the v3 cli dist binary with a graceful fallback if
the build skipped — keeps the job informational, not blocking.

Affected run: https://github.com/ruvnet/ruflo/actions/runs/25648092044
Triggered by main push of 31c6f97891 (alpha.15 federation publish).

Co-Authored-By: RuFlo <ruv@ruv.net>

* fix(agentdb): #1889 — symmetric memory-store fallback in pattern-search + paired-tool CI guard

## The bug

agentdb_pattern-store had a graceful fallback to memory_store namespace
'pattern' when the ReasoningBank controller is unavailable. agentdb_pattern-search
did NOT — it queried ReasoningBank only and returned 'controller: unavailable'
with empty results even when the pattern was sitting in the memory_store
'pattern' namespace from the store-side fallback.

The reporter (#1889) stored 8 patterns successfully and got zero results
from every search. The two MCP tools worked in isolation but their shared
contract (store → search round-trip) was broken because they read from
different substrates.

## The fix

agentdb_pattern-search gains a tiered symmetric fallback when bridgeSearchPatterns
returns null OR empty results:
  Tier 1: semantic search via searchEntries against namespace='pattern'
  Tier 2: substring scan via listEntries against namespace='pattern'
          (catches just-written entries before HNSW indexes them)

Both ends now report controller='memory-store-fallback' when the
ReasoningBank-side path is unavailable, so the round-trip is observable
and the symmetry is auditable.

## Answer to 'how did the CI guard miss this?'

It didn't catch it because every existing smoke (memory-import, mcp-protocol,
plugin-hooks, tool-descriptions) tests SINGLE tools in isolation. None of
them tested PAIRED-tool round-trip contracts — exactly the gap the bug
sits in. New scripts/test-mcp-roundtrips.mjs explicitly tests paired
MCP tools (store-then-search-by-sentinel) and a static dist-scan that
asserts both store and search have the memory-store-fallback path
present. New v3-ci.yml job 'mcp-roundtrip-smoke' wired into witness-verify
needs[].

## Witness manifest

103/103 fixes verified (was 102). New entry: #1889.

Closes #1889

Co-Authored-By: RuFlo <ruv@ruv.net>
EOF

* docs(plugins): add Verification & Discoverability section

Cross-references ADR-112 (tool description audit) + verification/ (witness
manifest) from the plugins README so users can find the guardrails that
gate plugin publishing.

Co-Authored-By: RuFlo <ruv@ruv.net>

* fix(ci): #1889 smoke — make behavioural round-trip advisory; dist-scan is the hard gate

The previous strict 'store-controller === search-controller' assertion
failed on CI because the bridge layer labels its own fallback writes
('bridge-fallback') differently from the memory-store-layer fallback
reads ('memory-store-fallback'). Different layers, different labels —
that's expected; the bug was the bridge writing AND search returning
0 results, not the label asymmetry.

Now the smoke:
  - HARD-fails if either dist-scan check is missing (memory-store-fallback
    path absent in store OR search). These are the load-bearing lines that
    prove the symmetric fallback exists in the shipped binary.
  - ADVISORY-notes the behavioural round-trip — if sentinel is found,
    pass; otherwise log the labels + result count and continue. Lets the
    smoke ship in CI environments where memory-db bootstrap is fragile.

Future work: align bridge and memory-store fallback labels OR have search
detect 'bridge wrote here' and read from the bridge's substrate too.

Co-Authored-By: RuFlo <ruv@ruv.net>

* fix(ci): #1889 smoke — dist-scan first, behavioral last with hard watchdog

Previous version ran behavioral round-trip before dist-scan. When the
memory backend bootstrap blocked the event loop, the inner setTimeout
never fired and the script never reached dist-scan — hanging the CI job
for ~10 minutes until GitHub killed it.

Restructured:
  Stage 1 (always runs): static dist-scan — store fallback present,
    search fallback present, both tool defs present. Exits non-zero
    if any check fails. Three load-bearing assertions; the durable
    contract for the #1889 fix.
  Stage 2 (advisory): behavioral round-trip with 25s inner-timeouts AND
    a 60s process-level watchdog. If anything hangs the watchdog SIGKILLs
    with a clean message and exit 0 — CI never hangs.

Local run confirms both stages log + script exits in <1s when memory
backend is happy; <60s in any pathological case.

Co-Authored-By: RuFlo <ruv@ruv.net>

---------

Co-authored-by: Reuven <cohen@ruv-mac-mini.local>
2026-05-11 00:29:58 -04:00

188 lines
16 KiB
JavaScript

#!/usr/bin/env node
/**
* ADR-112 Phase 1+2+3 — bulk applier for "Use when …" guidance suffixes.
*
* For each MCP tool name we know its category (agent / memory / agentdb /
* workflow / hooks / swarm / embeddings / claims / browser / cost /
* intelligence / aidefence / autopilot / federation / iot-cognitum / wasm /
* ruvllm / config / session / hive-mind / coordination / system / mcp /
* neural / progress / claims / transfer / daa / performance / analyze /
* guidance / ruvllm), and for each category we know the native-tool overlap
* and the Ruflo value-add. The script appends a category-appropriate
* "Use when … is wrong because …" suffix to any description that doesn't
* already include "Use when" / "Prefer over" / "Pair with" / "fall back".
*
* The script never modifies descriptions that already have guidance —
* agent_spawn, agent_execute, the seven we hand-wrote, etc.
*
* Run:
* node scripts/bulk-fix-tool-descriptions.mjs # dry-run
* node scripts/bulk-fix-tool-descriptions.mjs --write # apply
*/
import { readFileSync, readdirSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
const TOOLS_DIR = 'v3/@claude-flow/cli/src/mcp-tools';
const WRITE = process.argv.includes('--write');
const GUIDANCE_RE = /Use when|Prefer .* over|Pair with|fall back|native .* is (fine|wrong)/i;
// Category → "Use when …" suffix. Each suffix names the native overlap (if
// any) and the Ruflo value-add. Honest about when native is fine — that
// keeps Claude's trust in our guidance.
const SUFFIX = {
// -------- Memory & persistence --------
memory_: ' Use when native Read/Write is wrong because you need (a) cross-session retrieval by semantic similarity (vector embeddings) not by file path, (b) namespacing across projects without managing directory layout, or (c) the .swarm/memory.db audit trail. For one-shot file I/O, native Read/Write is fine.',
agentdb_: ' Use when generic memory_* tools are wrong because you need AgentDB-specific controllers (HNSW vector search, hierarchical tiers, causal-graph links, pattern store/recall, RaBitQ quantization). For simple key-value persistence, memory_store/memory_retrieve are simpler. For unrelated file work, native Read/Write are fine.',
embeddings_: ' Use when text similarity matters beyond keyword match — native Grep finds exact strings, embeddings find meaning. Pair with memory_store / agentdb_pattern-search to land the vector against your knowledge base. For literal symbol search, native Grep is faster.',
// -------- Agents & orchestration --------
agent_: ' Use when native Task is wrong because you need agent-lifecycle state (cost-tracking, taskCount across turns, swarm coordination, model routing via 3-tier). For one-shot subagents with no learning loop, native Task is fine.',
swarm_: ' Use when native Task tool is wrong because you need multi-agent coordination — topology (hierarchical/mesh/star), consensus (raft/byzantine/gossip/crdt/quorum), shared memory namespace, or anti-drift gates. For independent one-shot subagents, native Task is fine; spawn each separately.',
task_: ' Use when native TodoWrite is wrong because you need cross-session task persistence, agent assignment, dependency tracking, or completion analytics in the .swarm/memory.db. For in-session checklists native TodoWrite is simpler and faster.',
coordination_: ' Use when native Task is wrong because the work crosses multiple agents that need to vote/sync/load-balance — TodoWrite + a single Task cannot orchestrate consensus. For one-off subtask dispatch, native Task is fine.',
'hive-mind_': ' Use when native Task is wrong because you need queen-led collective intelligence — Byzantine-FT consensus, broadcast across many worker agents, shared memory with bounded conflict. For a single subagent, native Task is fine. Pair with swarm_init first to set topology.',
// -------- Hooks & lifecycle --------
hooks_: ' Use when native Bash hooks (via Claude Code\'s settings.json) are wrong because you need Ruflo-side state — pattern persistence, neural training signals, model-routing learning, cost tracking, audit chain. For one-off shell commands, plain Bash hooks are fine.',
// -------- Sessions --------
session_: ' Use when native conversation memory is wrong because you need durable cross-session state — restoring agent definitions, swarm topology, memory store, breaker history. For in-session continuation only, no tool needed. Pair with session_save before exiting and session_restore on resume.',
// -------- Config / system --------
config_: ' Use when native settings.json edits are wrong because the values need to be read by the Ruflo runtime (daemon, MCP server, neural router) — those load via the config_* path, not by re-reading settings.json. For .gitignore / .editorconfig style files, native Edit is fine.',
system_: ' Use when native Bash is wrong because you need Ruflo runtime metrics (HNSW index size, ReasoningBank state, swarm health, breaker status) — those are not in /proc, only in the running daemon. For OS-level info (uptime, disk, mem), native Bash + standard tools are fine.',
mcp_: ' Use when native Claude Code MCP status is wrong because you need Ruflo-side server detail — tool counts per namespace, transport stats, MCP handshake errors. For just "is MCP up?", `claude mcp list` is fine.',
status_: ' Use when generic Ruflo health checks are wrong because you want a single quick read of overall system state — daemon up?, swarm initialized?, memory db healthy?, federation peers connected? For deep debugging, prefer the dedicated tools each subsystem exposes.',
doctor_: ' Use when generic shell debugging is wrong because you want Ruflo-aware checks — Node/npm versions, daemon, memory DB, API keys, MCP servers, disk space. For unrelated environment troubleshooting, native shell + git/which/env are fine.',
// -------- Workflow --------
workflow_: ' Use when native TodoWrite + sequential Bash is wrong because the work has a real dependency graph that needs persistence, retry policy, pause/resume, and step-output binding across LLM-driven steps. For a single linear todo list, native TodoWrite is fine.',
// -------- Browser --------
browser_: ' Use when native WebFetch is wrong because you need real browser automation — JS-heavy SPA scraping, login flows with cookie reuse, replay against DOM-drifted versions, AIDefence PII gating before content reaches Claude. For static HTML pages, native WebFetch is faster and free.',
// -------- Security & defense --------
aidefence_: ' Use when nothing native exists — Claude Code does not have a PII / prompt-injection / adversarial-text scanner. Pair with any tool that ingests untrusted input (browser scrape, federation envelope, memory_import_claude).',
security_: ' Use when native package-audit (`npm audit`) is wrong because you need Ruflo-aware checks — known-bad dep patterns, secret detection, path-traversal in MCP inputs, witness chain verify. For just listing CVEs in your lockfile, native `npm audit` is fine.',
// -------- Federation --------
federation_: ' Use when nothing native covers cross-installation agent communication — Claude Code talks to its own MCP server only. Pair with federation_init first; once peers are joined, federation_send routes signed envelopes with PII gating, breaker, and audit. For local-only work, no federation tool is needed.',
// -------- Cost tracking --------
cost_: ' Use when native usage estimates are wrong because you need per-agent / per-model / per-task attribution across turns and sessions. The cost-tracking namespace persists between calls; reading the Claude CLI\'s built-in usage shows only the current turn. For one-shot cost checks, the native CLI suffices.',
// -------- Intelligence / neural --------
intelligence_: ' Use when native Task / Read prompting is wrong because you want learned-pattern routing — Ruflo\'s SONA neural router picks tier (Agent Booster / Haiku / Sonnet+Opus) based on past success on similar tasks. Pair with hooks_post-task to feed back outcomes. For one-shot prompts without learning, native Task is fine.',
neural_: ' Use when nothing native trains on your workflow — Claude Code has no learning loop. Use to train SONA/MoE/EWC patterns from successful task outcomes; query via neural_predict before spawning agents. Off-path for one-shot work.',
// -------- Autopilot --------
autopilot_: ' Use when running long-horizon goals that should resume automatically across sessions — Claude Code has no native autonomous-loop scheduler. Pair with autopilot_enable + a goal description, then let cron fires advance the work. For interactive single-task sessions, native Task is fine.',
// -------- DAA --------
daa_: ' Use when native Task is wrong because you need agents that adapt their cognitive pattern (convergent / divergent / lateral / systems / critical) per-task and share knowledge across the swarm. For static one-shot agents, native Task is fine.',
// -------- WASM agents --------
wasm_: ' Use when native Task is wrong because the workload needs sandboxed isolation — untrusted code execution, browser-side run, deterministic replay. Pair with wasm_gallery_search to find a published agent, or wasm_agent_create to scaffold a fresh one. For trusted in-process work, native Task is fine.',
// -------- RuVLLM (local inference) --------
ruvllm_: ' Use when sending every prompt to the Anthropic API is wrong because you need local inference — air-gapped environments, MicroLoRA-fine-tuned per-task adapters, or sub-cent per-call cost. For general Claude work native Task is the right call.',
// -------- Performance --------
performance_: ' Use when native shell timing (`time`, `hyperfine`) is wrong because you want Ruflo-aware benchmarks — HNSW search latency, breaker decisions/sec, MCP response p50/p95, embeddings throughput. For OS-level process profiling, native shell + perf are fine.',
perf_: ' Use when native shell timing is wrong because you want Ruflo-aware benchmarks (HNSW, swarm, MCP). For OS-level process profiling, native shell + perf are fine.',
benchmark_: ' Use when native `time`/`hyperfine` is wrong because you want a Ruflo-aware suite — agent latency, memory recall accuracy, neural routing hit rate. For OS-level micro-benchmarks, native shell is fine.',
profile_: ' Use when native Node `--prof` is wrong because you want Ruflo-component-specific traces (controller-by-controller, hook-by-hook, agent-by-agent). For low-level CPU/heap profiling, native Node profiler + clinic.js are fine.',
// -------- Analyze --------
analyze_: ' Use when native `git diff` / `grep` / static analysis is wrong because you want LLM-graded change classification, reviewer recommendations, or risk scoring. For literal-text inspection, native tools are fine.',
// -------- Progress tracking --------
progress_: ' Use when native TodoWrite is wrong because you need cross-session goal-completion tracking with witness/audit trail. For in-session checklists, native TodoWrite is simpler.',
// -------- Transfer / IPFS --------
transfer_: ' Use when native package install (`npm i`, `pip install`) is wrong because the artifact lives on IPFS (plugins, witness chains, learned patterns). For npm-registry deps, native npm is fine.',
// -------- Guidance --------
guidance_: ' Use when generic "what tool should I use?" guessing is wrong — Ruflo\'s guidance system uses the live tool index + your workflow context to recommend. Pair with hooks_route at task start. For trivial native-only tasks, no guidance call is needed.',
// -------- IoT (Cognitum Seed) --------
iot_: ' Use when native ssh-into-device is wrong because you need Ruflo-tracked fleet state — trust scoring, telemetry anomaly detection, witness chain verification. For one-off device debugging, native ssh is fine.',
// -------- Claims (authorization) --------
claims_: ' Use when nothing native covers per-agent capability gating — Claude Code agents have file-system access by default. Pair claims_grant + claims_check before letting an agent run privileged ops. For trusted in-session work, no claims call is needed.',
// -------- Terminal --------
terminal_: ' Use when native Bash is wrong because you need a persistent terminal session across turns/agents with output capture and replay. For one-shot shell commands, native Bash is fine.',
// -------- Daemon --------
daemon_: ' Use when native systemd/launchd is wrong because you want to manage just the Ruflo background workers (12 worker types, priority-aware) without touching OS-level service management. For OS-level service mgmt, native tools are fine.',
// -------- AgentDB causal/graph --------
causal_: ' Use when native bug tracker / postmortem doc is wrong because you want machine-readable cause→effect links queryable via Cypher. For human-readable postmortems, native markdown is fine.',
graph_: ' Use when native grep across files is wrong because you want typed entity-relation traversal — \"all decisions related to ADR-097\", \"all peers signed by this Ed25519 key\". For literal text search, native Grep is faster.',
// -------- ReasoningBank / search --------
reasoningbank_: ' Use when native Task is wrong because you want learned-trajectory replay — past successful approaches retrieved by current-task similarity. Pair with reasoningbank_judge + reasoningbank_distill to close the learning loop. For one-shot work without learning, native Task is fine.',
search_: ' Use when native Grep is wrong because you want semantic match (vector / hybrid / MMR-reranked). For exact-token search, native Grep is faster and free.',
};
const CATCHALL = ' Use when native Bash / file tools are wrong because this MCP tool exposes Ruflo-specific state or controllers that have no shell equivalent. For tasks that fit a one-line native command, prefer that.';
function suffixFor(name) {
for (const [prefix, suffix] of Object.entries(SUFFIX)) {
if (name.startsWith(prefix)) return suffix;
}
return CATCHALL;
}
let totalChanged = 0;
let totalSkipped = 0;
const perFile = {};
for (const f of readdirSync(TOOLS_DIR).filter(n => n.endsWith('.ts') && !n.endsWith('.test.ts'))) {
const filePath = join(TOOLS_DIR, f);
let src = readFileSync(filePath, 'utf-8');
let fileChanged = 0;
let fileSkipped = 0;
// Match `name: '...',` followed by `description: '...'` (with escaped chars)
// and replace the description with description + suffix if no guidance.
src = src.replace(
/(name:\s*'([^']+)',\s*\n(?:\s*[^,\n]+,\s*\n)?\s*description:\s*')((?:[^'\\]|\\.)*)(')/g,
(full, before, name, desc, close) => {
if (GUIDANCE_RE.test(desc)) {
fileSkipped++;
return full;
}
const suffix = suffixFor(name);
// Need to JS-escape any single-quote in suffix (template is single-quote literal)
const safeSuffix = suffix.replace(/'/g, "\\'");
const newDesc = desc.replace(/\s+$/, '') + safeSuffix;
fileChanged++;
return `${before}${newDesc}${close}`;
},
);
if (fileChanged > 0 && WRITE) {
writeFileSync(filePath, src);
}
perFile[f] = { changed: fileChanged, skipped: fileSkipped };
totalChanged += fileChanged;
totalSkipped += fileSkipped;
}
console.log(`Bulk tool-description fix (ADR-112)`);
console.log(`====================================`);
console.log(`Mode: ${WRITE ? 'WRITE' : 'dry-run'}`);
console.log(`Total descriptions updated: ${totalChanged}`);
console.log(`Total skipped (already had guidance): ${totalSkipped}`);
console.log(`\nPer-file:`);
for (const [f, s] of Object.entries(perFile)) {
if (s.changed > 0 || s.skipped > 0) {
console.log(` ${f}: +${s.changed} updated, ${s.skipped} kept`);
}
}
console.log(`\nRun with --write to apply. Re-run scripts/audit-tool-descriptions.mjs after.`);