232 Commits
Author SHA1 Message Date
ruv 8761c96a27 fix(ci): warm up the AgentDB registry before ADR-130 P3's timed assertions
graph-trajectory-hooks-smoke failed 3/3 times in CI on unmodified #3467
code (reverted my own unrelated attempt at a fix in the process — see
prior commit). Root cause identified: bridgeRecordFeedback()'s first-ever
call opens the AgentDB ControllerRegistry (new ControllerRegistry().
initialize({ embeddingModel: ... })), which resolves/attempts the ONNX
embedding pipeline before falling back to mock embeddings — a real,
one-time-per-process bootstrap cost. TEST 1 (trajectory-step on a
never-started trajectory) never reaches that code path, so TEST 2
(post-task) was the first and only call to pay it, entirely inside its
own <200ms budget: ~1000-1250ms on GitHub Actions' cold filesystem
cache, 20-46ms locally with a warm one. Added an explicit, untimed
warm-up call before TEST 1 begins so the one-time cost is paid during
setup, matching how a real long-lived Ruflo session actually behaves
(bootstrap once at startup, not per call). Verified 3/3 clean local
runs post-fix, TEST 1 now measuring 1-2ms (down from paying the same
hidden cost intermittently).
2026-09-27 15:52:07 -04:00
ruv 0d0ebbe305 fix(ci): add two new node:test-only ADR files to the vitest ratchet baseline
#3432 and #3437 add plugins/ruflo-adr/scripts/__tests__/import-root-exit-3097.test.mjs
and verify-read-safety-3147.test.mjs, both node:test-only (unlike
parser-relations-3096.test.mjs's dual-mode shim). Root vitest can't discover a
test suite in either, which the CI ratchet correctly flagged as 2 new
unexpected failing files. Same documented class as the plugin's four other
baselined test files (see the file's own comment on plugins/ruflo-x-gateway).
2026-09-27 13:40:58 -04:00
Rudy Celekli 24a80bd75d test(ci): isolate witness fixture and search mocks 2026-09-27 10:30:47 -04:00
ruv 9f3ac3b801 Merge branch 'pr-3412' into integ/train-3.45.1 2026-09-26 11:32:48 -04:00
ruv 4a6e063de7 Merge branch 'pr-3391' into integ/train-3.45.1 2026-09-26 11:32:47 -04:00
ruv 72a39572f2 Merge branch 'pr-3383' into integ/train-3.45.1 2026-09-26 11:32:47 -04:00
ruv 2f192d4eb9 Merge branch 'pr-3367' into integ/train-3.45.1 2026-09-26 11:32:47 -04:00
ruv 227af0b969 Merge branch 'pr-3424' into integ/train-3.45.1 2026-09-26 11:32:47 -04:00
rUvandClaude 4c2c62df27 dream(intelligence, evaluated): wire CLAUDE_FLOW_PRIOR_DECAY env override for ModelRouter's dormant decay primitive (#3349) (#3350)
* dream(intelligence): #3349 wire CLAUDE_FLOW_PRIOR_DECAY env override for ModelRouter's dormant decay primitive

ModelRouter's discounted-Thompson-sampling priorDecay (built/tested/
benchmarked in #3049) shipped permanently inert: DEFAULT_CONFIG.priorDecay
was hardcoded to 1 (disabled) with no env/config knob, unlike its sibling
maxUncertainty (envMaxUncertainty()). Adds envPriorDecay() mirroring that
exact pattern; default behavior is unchanged when the env var is unset.

Evaluation: stash-isolated discriminating test (1/3 new tests fails against
reverted source, 16/16 pass restored); full @claude-flow/cli suite and
tsc --noEmit both byte-identical baseline vs candidate outside the 3 new
tests; benchmark re-run reproduces the 2026-08-17 receipt exactly. Real
non-stationary recovery win in the 'low' bucket; a small but statistically
real stationary-accuracy cost in the 'med' bucket, accepted only under the
pre-existing ±1pp tolerance band from the original receipt — disclosed in
the issue/gist rather than oversold. Independent adversarial critic verdict:
CONFIRMED-WITH-CAVEATS.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_011tXcJWNao4vh57f4uFXutg

* docs(dream-cycle): append 2026-09-17 ledger row + backfill 09-13..09-16 gap/run verification

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_011tXcJWNao4vh57f4uFXutg

* fix(audit): register CLAUDE_FLOW_PRIOR_DECAY as a documented env-var escape hatch

scripts/audit-env-var-precedence.mjs (ADR-125/ADR-130) flagged the new
CLAUDE_FLOW_PRIOR_DECAY read in model-router.ts as undocumented CLI-flag
precedence. It's the same operator-knob shape as the adjacent, already-
registered CLAUDE_FLOW_MAX_UNCERTAINTY (envMaxUncertainty() is the pattern
envPriorDecay() mirrors) — no CLI invocation owns the router's persisted
lifetime, so there's no flag to wire precedence against. Registered
alongside it with the same justification.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_011tXcJWNao4vh57f4uFXutg

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-26 10:08:14 -05:00
Rudy Celekli 0702af8158 fix(witness): fail marker-preserving hash drift in strict mode 2026-09-25 14:03:38 -04:00
rUv abf6da6beb test(security): distinguish direct argv from shell splitting 2026-09-24 06:26:44 +00:00
rUv 2950997f3e test(security): exercise every shipped github-safe helper 2026-09-24 06:23:05 +00:00
rUv 82dac70fd7 feat(router): ADR-389–391 — refresh router helper, opt-in MiniLM embedder, benchmark gate (no default change) (#3410)
* docs(adr): ADR-389..391 — router: refresh the helper, use the real embedder, change defaults only on a measured win

ADR-389: add router.js to CRITICAL_HELPERS + re-sign, so installed copies get
the word-boundary fix (#3401); today they are frozen at first write.
ADR-390: hooks_route embeds patterns and tasks with the local MiniLM chain
(generateLocalEmbedding), hash embedder as the reported fallback.
ADR-391: frozen, labelled 150-200 prompt corpus; A/B/C/D (current, MiniLM,
typesafe-hash, typesafe-onnx); ADR-150-style AND-gate for promotion.

Co-Authored-By: RuFlo <ruv@ruv.net>

* feat(init): refresh router.js as a critical helper on upgrade (ADR-389, #3401)

Installed copies of .claude/helpers/router.js never received the #2257
word-boundary fix: helper-refresh only re-copies CRITICAL_HELPERS and
router.js was not on the list, so an April substring router kept routing
"sync and review latest issues" to tester.

- helper-refresh.ts: add router.js to CRITICAL_HELPERS; generator
  fallback also emits it (generateAgentRouter).
- sign-helpers.mjs / verify-helpers.mjs / executor.ts executeUpgrade /
  smoke-helper-signing-security.mjs: keep their copies of the list in sync.
- New test: stale April router.js is replaced by the package copy via the
  signed refresh path (throwaway key), after which the prompt no longer
  routes to tester.

helpers.manifest.json is NOT re-signed here. Until it is, verify-helpers
fails ("manifest has no entry for router.js") and the production refresh
is fail-closed blocked for all helpers.

Co-Authored-By: RuFlo <ruv@ruv.net>

* bench(router): ADR-391 labelled evaluation corpus (197 prompts, frozen)

Blind-labelled corpus for comparing routing candidates: 11 labels incl. none,
64 adversarial trap cases, deterministic stratified dev/test split (i%5 in
{0,2} -> dev), dependency-free validator that recomputes the split and prints
the sha256. Frozen at sha256 b0c1923b2813b61907304dff56c53d199b709797bfb9600a7c1f496cf0c0bf04.

Co-Authored-By: RuFlo <ruv@ruv.net>

* feat(hooks): ADR-390 — semantic router can use the real MiniLM embedder

hooks_route compared tasks with pattern keywords using a character hash
(generateSimpleEmbedding), which measures spelling, not meaning. This adds a
router-embedder abstraction (src/ruvector/router-embedder.ts):

- embedForRouter(texts, kind) -> { vectors, embedder: 'minilm'|'hash', reason? }
- MiniLM path uses generateLocalEmbedding ONLY (never the bridge-first
  generateEmbedding, #2312). backend !== 'onnx', a throw, or a non-384-d
  vector degrades EVERY text in the call to the hash, with a reason.
- Selection: CLAUDE_FLOW_ROUTER_EMBEDDER=minilm|hash; DEFAULT_ROUTER_EMBEDDER
  stays 'hash' until ADR-391's benchmark decides.

getSemanticRouter builds both the native VectorDb and pure-JS indexes from one
embedder, records which one actually built the index, embeds the query with
that same embedder, rebuilds when the requested embedder changes, and shares
one in-flight build between concurrent callers. If the query fails on MiniLM
after a MiniLM index was built, the index is rebuilt with the hash so the two
spaces never mix. hooks_route results gain `embedder` (+ `embedderReason` when
degraded); existing fields are unchanged.

routeTaskForBench(task, { embedder }) is an internal, bench-only entry that
runs the same local path as hooks_route (skipping the AgentDB pre-route and
the opt-in typesafe wrapper) for ADR-391.

Co-Authored-By: RuFlo <ruv@ruv.net>

* chore(audit): register CLAUDE_FLOW_ROUTER_EMBEDDER as a known env escape hatch

ADR-390's router embedder selector is process-lifetime MCP router state, like
CLAUDE_FLOW_DISABLE_NATIVE_ROUTER; the ADR-391 bench passes the embedder
explicitly rather than through a CLI flag.

Co-Authored-By: RuFlo <ruv@ruv.net>

* bench(router): ADR-391 first run — no candidate promoted; fix typesafe-router test (e)

Benchmark (run-bench.mjs) on the frozen 197-prompt corpus
(sha256 b0c1923b…), test split n=113:
  A current (hash)   25.7% acc, p95 1.06 ms
  B MiniLM (ADR-390) 36.3% acc, p95 6.21 ms   (+10.6 pts, +485% p95)
  C typesafe hash    26.5% acc, p95 1.20 ms
  D typesafe onnx    29.2% acc, p95 8.35 ms   (raw pick 44.3%, gated to 5.6%)
No candidate passes the ADR-391 AND-gate (B fails the relative latency
bound), so DEFAULT_ROUTER_EMBEDDER stays 'hash' and typesafe stays opt-in.
Receipt committed under benchmarks/router/results/. ADR-391 gains a Results
section with the findings (structural ceiling: no pattern returns
researcher/reviewer/none; latency criterion mis-specified for a ~1 ms base;
typesafe gate too tight) as owner decisions, not applied.

ADR-389 -> Implemented; ADR-390 -> Implemented as opt-in; ADR-391 -> Implemented.

typesafe-router test (e) was red on main since #3402 made the keyword
fallback word-bounded: it still asserted the old "latest"->tester fallback.
It now asserts the typesafe override and that a legacy fallback is kept.

Co-Authored-By: RuFlo <ruv@ruv.net>
2026-09-23 18:39:18 +00:00
rUv 72e17533ca fix(memory): release graph-edge-writer's native WAL handle so sql.js memory_store isn't refused (#3397) (#3405)
* fix(memory): release graph-edge-writer's native WAL handle so sql.js memory_store isn't refused (#3397)

graph-edge-writer cached its better-sqlite3 WAL handle in a module
singleton for the whole MCP server lifetime; nothing but the test-only
_resetBridgeDb ever closed it. Its -wal/-shm sidecars therefore stayed on
disk forever, and the #2735 sidecar guard (correctly) refused every later
sql.js whole-image write. On Windows, where the native AgentDB bridge is
off by default (#3024), that turned one hooks_post-task into a permanent
memory_store outage for the server.

- getBridgeDb() arms an unref'd idle timer (1s, CLAUDE_FLOW_GRAPH_EDGE_IDLE_MS)
  that checkpoints (TRUNCATE) and closes the handle; re-armed on every use.
- New releaseBridgeDb(dbPath?) is called by memory-initializer right before
  each #2735 guard, so a store immediately after an edge write works and the
  edge is checkpointed into the image sql.js rewrites.
- process 'exit' hook releases the handle so a clean shutdown does not leave
  stale sidecars for the next process. No signal handlers (would change
  Node's default termination).
- The guard itself is unchanged: an open WAL connection makes a whole-image
  write unsafe even in-process until its WAL is checkpointed.

Fixes #3397

Co-Authored-By: RuFlo <ruv@ruv.net>

* chore(audit): register CLAUDE_FLOW_GRAPH_EDGE_IDLE_MS as an env-only tuning knob (ADR-125)

Co-Authored-By: RuFlo <ruv@ruv.net>
2026-09-23 14:57:35 +00:00
ruvnetandClaude Opus 5 1116724a65 ci(guard): make the leaf-publish diff check actually run in PR CI
The guard's first real CI run passed but reported:

  NOTE: diff check NOT EXERCISED — could not diff against 'origin/main' (shallow clone?)

So only the range check ran; check 2 — the half that replays #3335/#3390 by
catching a leaf source change with no version bump — sat out entirely. The
NOT EXERCISED note is what made that visible instead of reading as a pass,
which is precisely why a skipped check must never print like a clean one.

Two causes, both fixed:

  - `git fetch --depth=1 origin $GITHUB_BASE_REF` fetches the base tip with no
    shared history, so `base...HEAD` dies with "fatal: no merge base". Now
    fetches --depth=50 with an explicit
    +refs/heads/<base>:refs/remotes/origin/<base> refspec.
  - The script only attempted three-dot. It now falls back to two-dot, which
    needs no common ancestor, and reports which mode actually ran.

Verified against the exact failing condition — an orphan branch with no merge
base against origin/main, where `git diff origin/main...HEAD` returns
"fatal: no merge base": the guard now reports "diff check ran (two-dot ...)"
rather than skipping. NC1 re-confirmed: a leaf source change without a version
bump still fails, now accompanied by "diff check ran (three-dot ...)" proving
the check executed rather than being skipped into a pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:20:11 -04:00
ruvnetandClaude Opus 5 f02759cd6d ci(guard): fail when a non-bundled @claude-flow/* leaf changes without a publish
Adds scripts/audit-leaf-package-publish.mjs, wired into v3-ci.yml's existing
audit job (semver is already installed there).

Only four @claude-flow/* components are bundled into the release tarballs —
INTERNAL_RUNTIME_PACKAGES in stage-internal-runtime-bundles.mjs: security,
codex, mcp, plugin-agent-federation. cli-core, neural, shared and memory
resolve from the REGISTRY at whatever v3/@claude-flow/cli/package.json pins, so
a merged source change in one of those does NOT ship with the three-package
train: the release silently carries the last published copy.

This has bitten twice. #3334 fixed RUFLO_INTELLIGENCE_MODE in
@claude-flow/memory; v3.42.1 shipped the train only and a fresh consumer
install had zero references to the fix (#3335). #3390's reasoningBank
activation spanned cli/src/memory/memory-bridge.ts AND
memory/src/controller-registry.ts — publishing only the train would have
shipped the bridge half and left the embedder half unreachable, leaving
reasoningBank exactly as inert as before. It was caught by hand and
memory@3.0.0-alpha.25 went out alongside v3.42.5.

What makes it easy to miss is that CI goes GREEN: the leaf's own package tests
run against workspace source, so the change looks verified while being
unshippable.

The audit asserts:
  1. every non-bundled leaf's workspace version is covered by the CLI's
     declared range (a miss means the release resolves to OLDER code than
     this repo's source);
  2. diff mode — if a leaf's publishable source changed vs the base ref but
     its version did not, fail with remediation.

It imports INTERNAL_RUNTIME_PACKAGES rather than restating it, so the bundled
and non-bundled sets cannot drift apart; exits non-zero if it inspects zero
leaves, so a broken classifier can't report clean; and reports a skipped diff
check as NOT EXERCISED rather than as a pass.

@claude-flow/shared is entered as a time-boxed accepted finding (expires
2026-10-21): the CLI pins 3.0.0-alpha.7 while the workspace is at
3.0.0-alpha.8, which IS published — stale, not unshippable. Raising the pin
edits v3/@claude-flow/cli deps and so forces a pnpm-lock.yaml regen in the same
change (#2540 -> #2552), which wants its own PR. An EXPIRED entry fails the
guard rather than muting it.

Negative-controlled, all four paths:
  - memory source changed, no bump      -> FAIL (replays #3390)
  - same change with a version bump     -> pass
  - test-only change in a leaf          -> pass (not publishable)
  - waiver expiry backdated             -> FAIL, "no longer honoured"

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 16:49:17 -04:00
Mr. MakandClaude Opus 5 b02d936935 fix(ci): print the failing assertion, not just the filename, on a ratchet failure (#3208)
The ratchet runs vitest with `--reporter=json --outputFile=…`, so nothing
vitest prints reaches the job log and the only record of an unexpected
failure is the file name. Recovering the assertion means downloading the
`test-results-ubuntu-latest` artifact — and a rerun replaces it, so after a
rerun attempt 1's assertion is gone for good.

The failure messages are already in the report this process parsed. Printing
them costs nothing, needs no new artifact, and survives a rerun, because job
logs are per-attempt.

`formatUnexpectedFailures()` is exported and pure so it unit-tests alongside
`evaluateTestReport()`. Every field it reads is optional by design: a
file-level abort ("No test suite found in file") carries `message` and no
assertions, and a report from another reporter version may carry neither. In
both cases the filename still prints, exactly as before, and the PASS path is
untouched.

Verified against a real failing run rather than a fixture: replaying run
35521170139's `vitest.json` through this script turns

      + v3/@claude-flow/cli/__tests__/memory-search-scores-3327.test.ts

into that same line plus

      ✗ memory_search score contract (#3327 Finding B) preserves standard fallback on import failure
        AssertionError: expected [ { key: 'a', …(5) } ] to deeply equal [ { key: 'a', …(4) } ]

which is the line that would have identified the failure in #3208 without a
second run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KTedGmPKpEdArVBMMn3JaP
2026-09-20 13:49:32 -04:00
Mr. MakandClaude Opus 5 0a2df2bfd2 fix(metaharness): use the @metaharness/* and ruflo CLI already installed instead of npm-installing at tool-call time (#3366)
On a stock ruflo 3.42.4 install every metaharness_* MCP tool reached for the
npm registry at tool-call time, although the packages it needs are already on
disk. Two causes, in the same six helper files, which is why they ship
together: the helper PINS rejected what ruflo installs, and three code paths
never looked for an installed copy at all. Fixing only the pins leaves darwin
and the memory tools on npx; fixing only the lookups leaves them resolving a
pin that nothing on disk satisfies.

1. Pins (metaharness, @metaharness/darwin)

The metaharness_* tools resolve MetaHarness through the plugin helpers' own
tilde pins, not through package.json. Those pins were never bumped:
_harness.mjs stayed at ~0.3.0 after the CLI declared metaharness ^0.4.1
(b8ec03f34), and _darwin.mjs stayed at ~0.8.0 while the CLI moved
@metaharness/darwin to ~0.9.0 (#2958) and ~0.10.2 (#3262). On a stock install
(metaharness 0.4.2, darwin 0.10.2) findLocalPackageDir() therefore rejected
the bundled metaharness, and every score / genome / mcp_scan / threat_model /
oia_audit / learn call ran
`npm install --prefix ~/.ruflo/metaharness-cache-0.3.0 metaharness@~0.3.0`
at tool-call time — a download of the older 0.3.2 (180 s install budget inside
a 120 s MCP subprocess budget), degraded offline. The darwin tools fetched
darwin 0.8.x through npx while gepa imported the bundled 0.10.2.

- _harness.mjs: METAHARNESS_PIN_VERSION ~0.3.0 -> ~0.4.1 (= CLI ^0.4.1)
- _darwin.mjs:  DARWIN_PIN_VERSION      ~0.8.0 -> ~0.10.2 (= CLI ~0.10.2)
- scripts/check-metaharness-pins.mjs: the watcher compared only the CLI's
  declared ranges (and MH_DARWIN_PIN), so it reported "all current" while the
  pins that decide what runs were two minors behind. It now reads each
  helper's *_PIN_VERSION: the pin must be a `~X.Y.Z` range (the only form
  _invoke.satisfiesTildeRange accepts), and must cover the declared range, or
  admit npm latest when the CLI declares nothing (redblue). An unreadable
  constant is fatal. `satisfies` now shares a `bounds()` helper with the new
  `covers()`; the results for the existing rows are unchanged.
- metaharness-pin-drift.yml also triggers on the three helper files.
- The two current-pin comments in TypeScript (distill-oracle.ts MH_DARWIN_PIN,
  metaharness-tools.ts ADR-153 header) named the old ~0.8.0 plugin pin; they
  now name ~0.10.2. Comments only, no behaviour.

2. No npm/npx at tool-call time (darwin, redblue, memory)

- _darwin.mjs ran `npx -y -p @metaharness/darwin@<pin> metaharness-darwin` on
  every evolve / bench / security_bench call: npm resolved the range each
  time, it ignored the optionalDependency the CLI ships, and it could run a
  different darwin than the gepa import in the same plugin. It now resolves
  like _harness.mjs — installed copy satisfying the pin, then the one-time
  versioned darwin-cache-<pin> that gepa already shares, then
  `process.execPath <abs bin>` with shell:false.
- _redblue.mjs always npm-installed into redblue-cache-<pin> on first use,
  never looking at the installed copy (ruflo 3.42.4 carries 0.1.4). It now
  checks the installed copy first. The chosen CLI path is realpath'd, because
  redblue only dispatches when import.meta.url === file://argv[1]; a
  pnpm-symlinked package dir (or a cache base under /tmp on macOS) otherwise
  makes it exit 0 having done nothing. Its spawn moves to process.execPath
  with shell:false: the argv carries MCP-supplied --config/--out/--in, which
  under the previous `shell: win32` went through cmd.exe unquoted.
- audit-list / audit-trend / oia-audit / similarity ran
  `npx @claude-flow/cli@latest memory ...` — an npm-registry resolution per
  call (the check _harness.mjs already removed for its own npx path), and
  possibly a different CLI version than the server that spawned the tool
  (#3306; cold-cache npx cost #3145/#3154). Without registry access oia_audit
  reported persisted:false and audit_list silently showed 0 records. They now
  call _invoke.runRufloCli(): process.execPath + bin/cli.js of the
  @claude-flow/cli that ships the plugin (<cli>/plugins/ruflo-metaharness, or
  v3/@claude-flow/cli in a checkout). A candidate needs dist/src/index.js next
  to bin/cli.js (mcp-launch.cjs's unbuilt-checkout guard) plus a package.json
  named @claude-flow/cli (a new check, so an unrelated parent directory can
  never qualify). CLI_CORE=1 still opts into `npx @claude-flow/cli-core@alpha`,
  and a layout with no usable local CLI still falls back to
  `npx @claude-flow/cli@latest`.
- drift-from-history.mjs: drop its unused CLI_PKG constant.

The two new lookups end in a spawn of a package bin with MCP-supplied argv, so
they pass findLocalPackageDir(..., { fromCwd: false }) and search only the
ruflo install that owns the plugin — never the tool's cwd or its ancestors.
The pre-existing metaharness lookup keeps its cwd walk unchanged.

Tests (both hermetic, published layout, stub packages, npm/npx PATH trap):

- test-pin-alignment.mjs (+ smoke 17z74a): 3 passed / 9 failed on main,
  12 / 0 here.
- test-no-tool-time-npx.mjs (+ smoke 17r2): 7 passed / 19 failed on main,
  26 / 0 here. It covers both darwin success branches (installed copy and the
  versioned cache), the memory round trip through the shipping CLI, the kept
  contracts (CLI_CORE, npx fallback, a bin/cli.js that is not
  @claude-flow/cli, darwin absent -> degraded exit 0), and the hardening
  above: a darwin/redblue in the tool's own cwd node_modules is never
  executed, a cache dir with package.json but no bin gets one reinstall
  instead of degrading forever, and neither helper spawns the resolved bin
  through a shell.

ADR-150 is unchanged: no static @metaharness import is added, no new
dependency, and every unresolvable path still returns the same
{degraded:true, reason} with exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KTedGmPKpEdArVBMMn3JaP
2026-09-20 11:55:51 -04:00
rUv 6f0ed71128 fix(release): spawn npm.cmd with shell:true in prepare-root-publish too (#3348)
Same EINVAL bug as #3346 (stage-internal-runtime-bundles.mjs's runBuild),
one file over: prepare-root-publish.mjs's own npm.cmd spawnSync for
building v3/@claude-flow/swarm and cli lacked shell:true, blocking the
root claude-flow package's Windows publish the same way. Found
immediately after #3346 while publishing 3.42.3 end to end.

Single-string shell:true form, matching #3346's fix, to avoid Node's
DEP0190 warning. Still constant literals, no injection surface.
2026-09-16 16:36:46 +00:00
rUv fcee45bc1d fix(release): spawn npm.cmd with shell:true on Windows (#3346)
* fix(release): spawn npm.cmd with shell:true on Windows

spawnSync('npm.cmd', ['run','build'], {stdio:'inherit'}) threw EINVAL on
Windows, blocking every @claude-flow/cli publish on this platform since
CreateProcess can't launch a .cmd directly and Node has refused to shell
out implicitly since CVE-2024-27980.

Found live while publishing 3.42.3: prepublishOnly -> prepare-publish.mjs
-> stageInternalRuntimeBundles() rebuilds the four internal runtime
packages (security/codex/mcp/plugin-agent-federation) and hit this on
the very first Windows publish attempt after upgrading past the Node
version where the implicit shell fallback still worked.

command/args here are constant literals ('npm.cmd', ['run','build']),
never hook-derived or user-controlled, so shell:true carries no
injection risk.
Co-Authored-By: RuFlo <ruv@ruv.net>

* fix(release): avoid DEP0190 warning in the npm.cmd shell:true fix

Single-string command form instead of args array — same shell:true
safety (still constant literals, nothing interpolated), no warning.
Co-Authored-By: RuFlo <ruv@ruv.net>
2026-09-16 16:11:14 +00:00
rUv 90889f4772 feat(sona): default the learning mode from RUFLO_INTELLIGENCE_MODE (#3334)
* feat(sona): default the learning mode from RUFLO_INTELLIGENCE_MODE

The SONA learning mode was only settable per call (`config.mode`), defaulting
to `balanced` otherwise. A deployment that wants every session to learn in a
specific profile — a managed desktop fleet running `research` for +55% quality,
or `edge` on constrained hosts — had no way to set that default short of
threading `mode` through every call site or patching a generated helper.

This reads `RUFLO_INTELLIGENCE_MODE` as the fleet-wide default in both the SONA
config merge (`@claude-flow/integration` sona-adapter) and the memory learning
bridge (`@claude-flow/memory`). Precedence is unchanged where a caller is
explicit: `config.mode` (or `sonaMode`) still wins; the env only fills the
former default slot; and an unset or unrecognised value returns undefined so
the code falls through to `balanced` — a typo can never silently select a
profile nobody asked for. Validated against the mode enum in each package.

Co-Authored-By: claude-flow <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01MUrJdDHhQ91hwpUBQhV4fC

* fix(sona): resolve sonaMode per-instance in LearningBridge, not at module load

learning-bridge.ts captured sonaModeFromEnv() into a module-scope
DEFAULT_CONFIG constant, evaluated once on first import — so any
RUFLO_INTELLIGENCE_MODE set later in the same process (including
test setup) was permanently missed. sona-adapter.ts's equivalent
(mergeConfig()) already re-reads the env var on every call; this
brings LearningBridge's constructor in line with that behavior,
preserving the same precedence (explicit config.sonaMode > env > default).

Adds regression tests for both packages (sona-adapter.ts had none
previously) and registers RUFLO_INTELLIGENCE_MODE in
KNOWN_ESCAPE_HATCHES per the audit script's existing requirement for
operator-knob env vars outside its SCAN_ROOTS.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01SjxUvZnyBdwtx4Y1jKAqt6
2026-09-16 02:22:13 +00:00
rUvandClaude e036e35dd7 dream(memory): #3231 wire embedding near-dup detection into MemoryConsolidator.dedup() (evaluated, ACCEPT) (#3232)
* dream(memory): wire embedding near-duplicate detection into MemoryConsolidator.dedup()

MemoryConsolidator.dedup() only ever deduplicated entries via byte-exact
SHA-256 content hashing, even though every entry already carries a
computed .embedding and the adapter's HNSWIndex (already held in scope
for removal bookkeeping) is incrementally kept in sync. Paraphrases and
reformattings of the same memory were never caught.

Adds a second pass: for hash-pass survivors with an embedding, query the
already-populated HNSWIndex for cosine-similarity neighbors above a
configurable similarityThreshold (default 0.95, matching the unwired
domain-layer consolidator's own constant), and merge using the same
keeper-selection strategies (now factored into shared selectKeeper/
mergeGroup helpers). Guarded to cosine-metric indexes only; disabled via
similarityThreshold >= 1.

Evaluation: 471/472 passing in @claude-flow/memory (1 pre-existing,
unrelated, environmental failure); baseline (git-stash-isolated source)
fails exactly the 2 new discriminating tests and passes the rest,
confirming the fix is additive and non-regressive. tsc --noEmit clean.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG

* dream(memory): fix-forward — converge near-dup pass to a fixed point in one dedup() call

Independent adversarial critique (STEP 10) reproduced a real, bounded gap
in the prior commit: a near-duplicate cluster larger than
NEAR_DUP_SEARCH_K (8) split into multiple leftover sub-group survivors
within a single dedup() call, because each round's `consumed` bookkeeping
permanently excluded a group's keeper from further matching even though
it was still fully present in the index. 15 pairwise-identical-embedding
entries collapsed to 2 survivors (merged: 13) instead of 1 (merged: 14).

It self-healed across repeated runAll() calls (background timer /
nightlyLearner), so this was never a permanent-data-loss bug, but a
single manual dedup() call under-converged relative to its documented
"collapse duplicates" contract.

Fix: loop pass 2 to a fixed point (re-scan until a full round produces
zero merges) instead of a single scan. Always terminates — entries
strictly decrease each round that merges anything. `groups` now counts
merge operations across all rounds, which can exceed the number of
underlying duplicate clusters when one needed more than one round;
documented in the method's doc comment.

Added a discriminating regression test reproducing the critic's exact
15-entry scenario; confirmed it fails against the pre-fix single-round
code (merged: 13) and passes against this fix (merged: 14, single
survivor). Full package suite: 472/473 passing (1 pre-existing,
unrelated, environmental failure, unchanged from before this commit).
tsc --noEmit clean.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG

* dream(memory): 2026-09-08 research gist — embedding near-dup consolidation

Full SOTA report: 5-role parallel research fan-out, ledger check with
GitHub-verified fates for the last 14 rows (3 merged since 09-03, one PR
now crossing the 14-day stale threshold for the first time this cycle),
competitor comparison (CrewAI/Mem0/Zep/LangMem/Qdrant/SemDedup), plugins
and automation scan findings, adversarial critique with a real bug found
and fixed, and witness stamp.

No gist-creation MCP tool available in this session (GitHub MCP tools
cover issues/PRs/repos, not gists) — committing to docs/dream-cycle/
instead, matching every dream-cycle night since 2026-08-14.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG

* dream(memory): #3231 update LEDGER.md for 2026-09-08

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG

* dream(memory): #3231 empirical receipts for dedup() near-dup pass (response to review)

Response to PR #3232's INCONCLUSIVE review, which asked for a threshold
contract rather than a smoke fixture: a contamination-checked embedding
corpus with pinned model, precision/recall including hard negatives and
cross-namespace cases, baseline vs. candidate latency/memory at realistic
cardinality, and deterministic replay + unchanged exact-dedup behavior.

Real pinned-model corpus: attempted directly against both ONNX backends
this monorepo ships. Both fail for reasons unrelated to this candidate,
reproduced on an unmodified checkout before writing anything:
- @huggingface/transformers@4.2.0: transformers.node.mjs does
  `import { Tensor } from "onnxruntime-common"` against a CJS module — a
  real ESM/CJS interop break under this sandbox's Node version.
- @xenova/transformers@2.17.0: pulls in sharp@0.32.6, whose native
  binding (sharp-linux-x64.node) isn't present for this platform — the
  same failure independently reproduced moments later by this package's
  own nightlyLearner test, which falls back to mock embeddings.
Neither is fixable within this PR's scope without touching unrelated,
pre-existing native/module-resolution infrastructure. Documented in
detail at the top of the new test file rather than silently worked
around or omitted.

What's delivered instead, in
v3/@claude-flow/memory/src/consolidator-embedding-benchmark.test.ts:
- A deterministic synthetic corpus where pairwise cosine similarity is
  constructed EXACTLY via vectorAtSimilarity() (Gram-Schmidt against a
  random orthogonal vector), not measured after the fact — a rigorous
  instrument for threshold mechanics specifically, disclosed as not
  validating real-world semantic accuracy.
- Precision/recall/false-merge-rate across a 5-point threshold sweep,
  every point at a deliberate margin from every corpus similarity value
  after an exact-boundary collision (sweep value == corpus value) proved
  float32-jitter-sensitive while writing this — documented as a finding,
  not silently patched around.
- Namespace isolation: proved cross-namespace near-dup merging matches
  cross-namespace hash-exact merging exactly (both pre-existing,
  unscoped-by-namespace behavior — this candidate doesn't change it).
- Determinism: 3 independent runs, compared by deterministic entry KEY
  (not raw id, which is randomly generated per store() and was a second
  bug this file's own first draft had to fix).
- Latency/memory at N=5000 (matching this repo's own HNSW-benchmark
  convention): initially measured 9.8s for the candidate pass, which
  traced to an unset `ef` search parameter defaulting to efConstruction
  (200) — every dedup() search traversed a 200-candidate list for a
  duplicate-detection task that only needs very-close neighbors. Fixed
  by passing an explicit NEAR_DUP_SEARCH_EF=32 in consolidator.ts,
  re-measured at 3.7s (~2.7x), precision/recall unchanged (still exact).
  Real numbers logged in the test output either way, not asserted away.

Full package suite: 478/479 passing (472 pre-existing + 7 new; same 1
pre-existing unrelated environmental failure as every prior commit on
this branch). tsc --noEmit clean.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG

* dream(memory): #3231 add new benchmark file to CI test-ratchet baseline

CI on a2e99cf failed with "1 unexpected failing file(s): +
v3/@claude-flow/memory/src/consolidator-embedding-benchmark.test.ts".

Root-caused via the vitest JSON report artifact (test-results-ubuntu-latest,
downloaded and parsed directly, not guessed): a suite-level collection
failure, "Cannot find package '@claude-flow/security' imported from
'.../v3/@claude-flow/memory/src/agentdb-retrieval-guard.ts'" — the same
unbuilt-sibling-package gap already accepted in scripts/ci-test-baseline.txt
for 9 other @claude-flow/memory test files, including consolidator.test.ts
itself (line 95). Every file in this package fails identically in this CI
checkout state; my new file just wasn't in the baseline yet because it's
new. Not a regression this PR introduced — confirmed by the failure being
a pre-import-time module-resolution error, not a test assertion in the
new file's own logic (which passes locally in a checkout with @claude-flow/
memory's dependencies actually built: see the 6/6 local run in the prior
commit's message).

Fix: add the new file to the baseline, alongside its 9 memory-package
siblings already there, matching this repo's own established convention
for this exact gap rather than working around it some other way.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG

* dream(memory): #3231 baseline 2 more federation/gateway plugin test files

CI on 20afdc8 failed (run 34905889339) after main was rebased in — 3
unexpected failing files, all pre-existing, unrelated to this PR's diff:

  plugins/ruflo-chatgpt-federation/test/oauth.test.mjs
    Cannot find package 'jose'
  plugins/ruflo-chatgpt-federation/test/publisher.test.mjs
    Cannot find package 'nostr-tools/pure'
  plugins/ruflo-x-gateway/test/oauth.test.mjs
    Cannot find package 'jose'

Verified via the vitest.json artifact (downloaded and parsed, not
guessed): all three are plugin-local-dependency collection failures,
identical in kind to the already-baselined sibling
plugins/ruflo-x-gateway/test/gateway.test.mjs ("node:test suite with its
own deps... root vitest cannot resolve them"). Confirmed these are new
on main itself (main's own ci-test-baseline.txt at the rebase point does
not list them either) and unrelated to v3/@claude-flow/memory — nothing
this PR touches.

Re-verified locally: re-ran the full @claude-flow/memory suite against
the rebased main (496/497 passing, same 1 pre-existing unrelated
failure as always) and recomputed the ratchet's unexpected-set directly
against the CI run's vitest.json artifact — 0 unexpected after this
change.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-14 23:13:14 +00:00
rUvandClaude 1bd734e272 dream(security): #3102 bind ADR-377 caller-identity verification into live authorizeMcpTool chokepoint (evaluated, ACCEPT-with-caveats) (#3103)
* dream(security): #TBD bind ADR-377 caller-identity verification into live authorizeMcpTool chokepoint (evaluated, ACCEPT-with-caveats)

authorizeMcpTool() trusted the plain, unsigned CLAUDE_FLOW_PRINCIPAL_ID env
var as caller identity with zero verification (OWASP ASI07-class gap).
ADR-377 Phase 3 already implemented Ed25519 issueInvocationToken/
verifyInvocationToken but never wired it into a live dispatch path.

Binds the two: DualModeOrchestrator mints a per-worker signed
InvocationToken at spawn (worker-lifetime TTL, wildcard tool scope --
disclosed reduction from the primitive's original per-call design);
authorizeMcpTool verifies it before trusting identity, failing closed on
missing/forged/expired tokens. Off by default (CLAUDE_FLOW_MCP_CALLER_AUTH).

6 new deterministic Vitest scenarios plus an independent adversarial
critique (fresh session, verdict CONFIRMED-SAFE-WITH-CAVEATS) are recorded
in docs/dream-cycle/dream-gist-2026-08-26.md. Darwin skipped (scope
mismatch -- binary crypto gate, no tunable parameter space).

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_015bfVASxkNXKpZbMSLhVU1L

* fix(ci): register ADR-377 Phase 3 caller-identity env vars as known escape hatches

CLAUDE_FLOW_MCP_CALLER_PUBKEY and CLAUDE_FLOW_MCP_INVOCATION_TOKEN
(DualModeOrchestrator's per-worker Ed25519 credential pair, read by
resolveMcpCallerIdentity()) weren't registered in
audit-env-var-precedence.mjs's KNOWN_ESCAPE_HATCHES, so the audit flagged
CLAUDE_FLOW_MCP_CALLER_PUBKEY as a real violation (exit 1). Same no-CLI-flag
reasoning as the existing CLAUDE_FLOW_PRINCIPAL_ID entry: these are
credentials minted into a spawned worker's own environment, not a value
any caller should be able to select via a flag.

Verified: audit-env-var-precedence.mjs now exits 0. policy-runtime.test.ts
still 18/18.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01RYP825phHMjPRmmjkcAWpq

* fix(ci): regenerate root package-lock.json for codex's new @claude-flow/security dep

The CI Test Suite job runs plain `npm ci` at repo root (no v3 pnpm
install step), so it resolves workspace deps strictly from
package-lock.json — not from v3/pnpm-lock.yaml, which this PR had
already updated. codex/package.json's new @claude-flow/security
dependency was missing from package-lock.json's codex entry, so CI's
npm ci never linked it and every dual-mode test importing
orchestrator.ts failed with "Cannot find package '@claude-flow/security'".

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01RYP825phHMjPRmmjkcAWpq

* fix(ci): build @claude-flow/security before Test Suite runs it

Complements d492ea835 (codex's package-lock.json entry) rather than
duplicating it: that commit fixed `npm ci` not *linking*
@claude-flow/security for codex at all ("Cannot find package"); this
fixes the fact that even where it IS linked (codex now, @claude-flow/cli
already), nothing builds its dist/ before tests run, so importers hit
ERR_MODULE_NOT_FOUND on dist/index.js instead.

This was already a known, pre-existing gap for @claude-flow/cli:
v3/@claude-flow/cli/__tests__/policy-runtime.test.ts has carried a
scripts/ci-test-baseline.txt entry for exactly this reason since before
this PR existed. dual-mode.test.ts and dual-mode-stdin-2947.test.ts
(added by the caller-identity candidate on this branch) newly trip the
same gap in @claude-flow/codex, confirmed by diffing main's own Test
Suite ratchet failures (3, pre-existing, unrelated oauth/publisher
plugin tests) against this branch's (5: the same 3 plus these 2). Adding
new files to ci-test-baseline.txt to paper over this is explicitly
forbidden by that file's own header ("Never add entries to make a
regression green"), and gutting the tests to avoid the import would
remove real coverage of the orchestrator/security integration point.

Real fix: build @claude-flow/security before "Run all tests", using its
own existing `npm run build` (tsc && verify-oauth-exports.mjs). Verified
locally end-to-end by removing both dist/ and the stale
tsconfig.tsbuildinfo (simulating a fresh CI checkout, since both are
gitignored) -- the build step produces a full dist/, both new codex test
files pass, and (bonus, explicitly allowed by the baseline file's own
"remove entries when fixed" rule) the pre-existing policy-runtime.test.ts
baseline entry is now consistently green too, so it's removed here.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_015bfVASxkNXKpZbMSLhVU1L

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-14 21:05:45 +00:00
rUv 8008bc3562 fix(federation): verified event identity must win over publisher-controlled content (#3322)
* fix(federation): verified event identity must win over publisher-controlled content

fetchRecent/fetchManyOn/fetchChannel spread the parsed message content
after the signature-verified id/pubkey/created_at, so a publisher could
put those keys in their own content and overwrite their verified
identity downstream (reduceClaims trusts .pubkey for claim
release/handoff authorization). Reorder so verified fields always win.

Also in x-federation-channels.ts: reqEvents() never called verifyEvent()
at all, so channel-read/grant-accept events were trusted unsigned; add
the check. Same content-override reorder in x_federation_channel_read.

fix(hooks): escape argv for the Windows shell:true npm-shim path

ruflo-hook.cjs's invokeHook() passes hook-derived values (command text,
file paths) through cmd.exe via shell:true to resolve npm's .cmd/.ps1
shims. Node does no escaping in that mode, so a value containing a
cmd.exe metacharacter could be reinterpreted as a separate command or
redirection instead of reaching the CLI as data. Add escapeCmdArg()
(qntm.org/cmd algorithm, same one cross-spawn uses) and apply it to
every argv element on that path.

No native Windows/cmd.exe available to execute this end-to-end in this
environment — verified at the string-transform level only (8 unit
tests). Real Windows execution remains a required follow-up gate.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01RYP825phHMjPRmmjkcAWpq

* fix(ci): register new hook test env var and test file with CI guards

Follow-up to the previous commit's escape-cmd-arg.test.cjs:
- Register RUFLO_HOOK_UNIT_TEST in audit-env-var-precedence.mjs's
  KNOWN_ESCAPE_HATCHES, matching the existing RUFLO_HOOK_CLI_OVERRIDE/
  RUFLO_HOOK_DEBUG_STDOUT entries for the same file.
- Register plugins/ruflo-core/scripts/escape-cmd-arg.test.cjs in
  ci-test-baseline.txt, matching the sibling mcp-launch.test.cjs entry —
  both are node:test-based .cjs files that vitest's own JSON reporter
  marks "failed" (it doesn't recognize the node:test API) even though
  every test inside genuinely passes under `node --test`.

Verified locally: node scripts/audit-env-var-precedence.mjs now passes
(exit 0). A scoped `vitest run` + `ci-test-ratchet.mjs --report` against
just these two files confirms escape-cmd-arg.test.cjs is no longer an
unexpected failure (the one remaining flagged file is a stale local
worktree artifact under .claude/worktrees/, not part of a clean CI
checkout).

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01RYP825phHMjPRmmjkcAWpq
2026-09-14 17:12:05 +00:00
rUv 8485302f10 feat(x-gateway): 0.6.1 — bound Seraphina by budget instead of an admin token (#3275)
* feat(x-gateway): 0.6.1 — Seraphina bounded by budget, not by an admin token

Squashes what is actually running on x.ruv.io so git and production agree.

Seraphina was gated alongside the write tools, but it is not like them: it reads
the roster, the claims board and recent messages, asks a model, and returns
advice. It writes nothing and carries no authority. The gate answered the wrong
question — the exposure is model spend, and spend is bounded with a budget, not
a password.

The practical cost was worse than a wrong abstraction. A browser cannot hold a
bearer secret, so a published UI was either locked out or tempted to ship
RUFLO_ADMIN_TOKEN to the client — the same token that mints invites and publishes
as the gateway identity. The safe-looking option was the catastrophic one.

Seraphina now answers with no token, bounded by a shared daily cap and a
per-client hourly cap, and anonymous callers cannot select the high or ultra
tiers. An admin token lifts both. Every write tool keeps its gate; verified live
that claims_issue, federation_invite_mint, federation_publish, federation_admit
and channel_publish all still refuse an anonymous caller.

Not included, deliberately: a tool to relay member-signed events through the
gateway. It was built, tested against the live relay, and removed. buzz-relay
refuses any EVENT whose pubkey differs from the NIP-42 authenticated connection,
so a gateway cannot publish on another identity's behalf — and it does not need
to. Members already publish as themselves through the wss://x.ruv.io proxy,
verified end to end. Shipping a tool that always fails would be worse than none.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* feat(x-gateway): serve onboarding guidance, and deliberately not key generation

The registry said "generate a Nostr keypair" and stopped, and it never named the
two identities in play. That gap was not theoretical: a capable agent inspected
the surface, found only gateway-signed publishing, and concluded ordinary members
could not publish from a dashboard at all. They can. It just was not written
anywhere they could reach.

federation_onboarding and ruv://federation/onboarding now answer it, open, no
token. The guide leads with the distinction that causes the confusion — your key
signs for you, the gateway's key signs for the service, and the gated tools are
gated precisely because they speak as the service. It carries the five steps, the
four things never to do, and the three traps this service has actually taught us:
sign the NIP-42 relay tag with the canonical URL even when proxied, the relay
binds publishing to the authenticated connection, and the channel tag is `c` not
`h`.

It does NOT generate keys, and that is the point rather than an omission. A
service that mints your keypair and returns the secret has seen your secret, and
becomes custodian of every identity it "helped" — the same custody mistake as
putting an admin token in a browser, inverted. The guide ships the code so the
caller runs it locally and the key never crosses the wire.

One test note worth keeping: the first version of the safety assertion regex-
matched for "send your secret" and failed on the guide's own "Never send your
secret key" line. It is now structural — no field may be named for secret
material — because a string search cannot tell an instruction from its negation.

Verified live on 0.7.0 (ruflo-x-gateway-00015-qj7) with an anonymous call. 20/20.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* fix(ci): register the gateway's credential and budget knobs, and skip test/ like tests/

The ADR-125 precedence audit failed the gateway PR, and it was right to look —
but for two reasons that are both about the audit, not the code.

RUFLO_ADMIN_TOKEN was not registered as an escape hatch even though its sibling
RUFLO_X_ADMIN_TOKEN is, with the same reasoning: a secret must never be a CLI
flag, where it lands in shell history and process listings. The gateway is a
long-running service with no typed command surface at all, so there is no
invocation to attach a flag to. Same for the two Seraphina budget knobs added in
#3275.

The other half was a naming gap. SKIP_DIRS already excludes `tests` and
`__tests__` because the audit is about production precedence, not test setup —
but the flagged lines were in `plugins/ruflo-x-gateway/test/`, singular, which
was not in the set. Two `process.env.X = 'test-admin-token'` assignments inside a
test fixture were reported as undeclared production reads.

Verified: the audit now exits clean on this branch.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
2026-09-10 19:20:21 +00:00
rUv dbf450a927 feat(federation): ruflo CLI + MCP integration for x.ruv.io, and Seraphina (swarm queen) (#3256)
* feat(federation): ruflo CLI + MCP integration for x.ruv.io, and Seraphina (swarm queen)

Integrate the open swarm federation into ruflo itself, and add Seraphina —
a primary-coordinator / swarm-queen guidance tool in the style of the ruOS
assistant terminal (invoked from a terminal or any MCP client).

MCP tools (src/mcp-tools/x-federation-tools.ts), registered in the barrel and
the in-process registry so `callMCPTool` and the MCP server both see them:
  x_federation_sync / roster / claims / registry   (open reads)
  x_federation_publish / invite_mint / admit       (gateway-identity writes,
    require RUFLO_X_ADMIN_TOKEN; fail closed with no network call)
Each description follows ADR-112 ("Use when … wrong because …").

Seraphina (src/mcp-tools/seraphina-tools.ts): `seraphina_guidance { goal }`
gathers the live roster, claims board and recent messages from the x.ruv.io
gateway, compacts them (dedupe by from|type, cap 15 — a cheap tier drowns in
repeated PeerHellos), and asks the cognitum meta-llm gateway
(https://api.cognitum.one/v1/messages, model cognitum-auto by default with a
tier override) using a queen system prompt that treats message content as
data, respects one-owner-per-resource claims, and returns
{ guidance, proposals[], risks[] }. Proposals are advisory. JSON is extracted
by slicing the outermost object so fenced/prose-wrapped answers still parse.
Key from SERAPHINA_METALLM_KEY (GCP secret seraphina-metallm-api-key).

CLI: `ruflo federation sync|roster|claims|registry|invite|admit|publish`
(src/commands/federation.ts), thin over the same tools.

Skill: .agents/skills/open-federation/SKILL.md documents CLI, tools,
Seraphina, the claims rules, the canonical NIP-42 relay-tag gotcha, and
onboarding.

Tests: 11 (x-federation: RPC shaping, SSE parse, resource mapping, fail-closed
admin gating, isError surfacing; seraphina: registration, fail-closed key,
context gathering + queen prompt + x-api-key, tier override, fenced-JSON
extraction). Project typechecks at 0 errors.

Verified live: Seraphina against the real 3-node federation returned 3
structured proposals + 4 risks via cognitum-auto (routed to a cheap tier).

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* fix(federation): ADR-125 env-var precedence — URL config takes a flag/arg, credentials registered env-only

The env-var-precedence audit correctly flagged the new process.env reads.
Config URLs now have a precedence path: `ruflo federation --gateway` feeds a
`gatewayUrl` tool arg (and Seraphina takes `metaLlmUrl`/`gatewayUrl`) that
wins over RUFLO_X_GATEWAY_URL / SERAPHINA_METALLM_URL, documented per ADR-125.
Credentials (RUFLO_X_ADMIN_TOKEN, SERAPHINA_METALLM_KEY) are registered as
env-only escape hatches with rationale: a secret must never be a CLI flag.
+1 test (arg precedence, arg not forwarded upstream). Audit passes locally.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* feat(federation): `ruflo federation join --code` — self-service invite access with your own key

The user-facing path into the open swarm. Decentralized by design: the user
generates/holds THEIR OWN Nostr key (~/.ruflo/nostr.key, 0600), redeems an
invite code with a NIP-98-signed claim directly against the relay (no admin in
the loop), proves membership via NIP-42, and can then publish as themselves.
The gateway never signs for a user.

- src/mcp-tools/x-federation-join.ts: x_federation_join { code, relayHttp?,
  relayWs?, keyFile? } — validates the code shape before any network call,
  loads/creates the key, NIP-98 claim, NIP-42 verify, returns pubkey + role.
  ADR-125: args take precedence over RUFLO_X_RELAY_HTTP / RUFLO_X_RELAY_WS /
  RUFLO_NOSTR_KEY_FILE.
- CLI: `ruflo federation join --code v2.…`
- nostr-tools added as an optionalDependency (secp256k1/Schnorr is not in
  node:crypto); the tool degrades with an install hint when absent. pnpm
  lockfile regenerated in this PR (frozen-lockfile CI).
- 4 tests: 0600 key create/reuse, valid NIP-98 header (kind 27235, verifies,
  bound to url+method+payload hash), malformed code rejected pre-network,
  ADR-112 description. Suite: 16/16; tsc 0; env-var audit passes.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* fix(federation): default relay -> wss://relay.ruv.io; join test tolerates missing nostr-tools

- x-federation-join defaults now point at the canonical relay.ruv.io host (Cloud Run
  domain mapping for buzz-relay); the raw run.app host stays routable for old clients.
- Validate the invite-code shape before the optional-dependency check so a bad code
  fails fast whether or not nostr-tools is installed.
- The root Test Suite runs the CLI tests via root npm ci, which never installs the
  CLI's optionalDependencies: the crypto cases are it.skipIf(!nostr-tools) so the
  file no longer trips the CI test ratchet.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
2026-09-10 02:51:40 +00:00
rUv 6396ff636c feat(x-gateway): open Nostr swarm federation gateway — MCP + ruv:// + ws proxy (deployed at x.ruv.io) (#3255)
* feat(x-gateway): MCP + ruv:// gateway for open Nostr swarm federation

New plugin `ruflo-x-gateway` — the service behind x.ruv.io. An MCP server
(Streamable HTTP at /mcp) that exposes ruflo swarm FEDERATION and CLAIMS over
an open, membership-gated, SIGNED Nostr relay (buzz-relay).

Why Nostr: every coordination message is a signed Nostr event, so authorship is
cryptographically verifiable and anyone the relay admits can participate — open
but secure. The relay gates membership + NIP-42 auth; no pre-pinning needed.

Surface:
- Tools: federation_identity / federation_join / federation_publish /
  federation_sync, claims_issue / claims_release / claims_status.
- Resources: ruv://federation/registry, ruv://swarm/roster, ruv://claims/board.
- src/nostr-federation.mjs: NIP-42 authenticated connect, signed publish, and
  verified fetch of #t=ruflo-swarm coordination events.
- src/server.mjs: node http server routing /, /health, /mcp (stateless
  Streamable HTTP), plus the ruv:// resources.
- Dockerfile for Cloud Run; persistent Nostr identity at /data (0600).

Smoke-tested locally: server starts, /health + / respond, POST /mcp tools/list
returns the tool set over SSE. Federation tools require relay membership (by
design) — deployment wires membership + DNS (x.ruv.io) as follow-up.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* feat(x-gateway): stable identity via RUFLO_NOSTR_KEY_HEX (GCP secret) + read-only-FS fallback

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* feat(x-gateway): v0.2.0 — ws proxy, admin-gated writes, invite/admit tools, tests

Security model made explicit: /mcp is public, so every tool that writes
using the GATEWAY's own identity (join, publish, claims_issue/release,
invite_mint, admit) now requires `adminToken`, checked constant-time and
fail-closed (no configured token => all writes denied). Reads and ruv://
resources stay open. Users publish with THEIR OWN keys via invite->claim.

- src/ws-proxy.mjs: transparent WebSocket proxy so wss://x.ruv.io (and
  /relay) fronts the Nostr relay; 500-conn cap, 502/503 on failure.
- src/relay-admin.mjs: mintInvite (NIP-98 POST /api/invites) and
  admitMember (NIP-43 kind 9030) — the gateway holds relay admin role, so
  the owner key never leaves GCP.
- src/security.mjs: per-IP token-bucket rate limit (60/min), 256KB body
  cap enforced before buffering, security headers, timingSafeEqual admin.
- src/claims.mjs: owner-per-resource reducer extracted for testing.
- src/server.mjs: createGateway() factory (testable), stateless MCP.
- test/gateway.test.mjs: 8 tests — claims rules, gating, rate limit, body
  cap, NIP-42 against a mock relay (verifies the signed challenge), routes.

Verified live before this commit: gateway pubkey admitted + promoted to
relay admin; invite minted and a fresh key self-joined via claim and
passed NIP-42 auth; deployed gateway publishes/reads over the relay.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* fix(x-gateway): surface canonicalRelay — relay verifies NIP-42 relay tag strictly

Empirical: auth through the wss://x.ruv.io proxy is ACCEPTED when the client
signs relay=<canonical relay URL> and REJECTED (verification failed) when it
signs relay=wss://x.ruv.io. Expose canonicalRelay + authNote at GET / and in
ruv://federation/registry so clients sign the right tag.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* ci: retrigger test-suite (ratchet flake; main green at ac08e4e02)

* feat(x-gateway): v0.3.0 — Seraphina swarm-queen guidance tool (admin-gated, cognitum meta-llm)

seraphina_guidance reads the live roster/claims/recent from the relay, compacts
them (dedupe by from|type, cap 15), and asks https://api.cognitum.one/v1/messages
with cognitum-auto (tier override). Admin-gated because it spends meta-llm
budget. JSON extracted by slicing the outermost object so fenced answers parse.
Key from SERAPHINA_METALLM_KEY. +1 test (9/9).

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* ci(ratchet): baseline ruflo-x-gateway node:test file (own deps; runs via plugin npm test)

Root vitest sweeps plugins/**/test/*.test.mjs and cannot resolve the gateway's
deps (nostr-tools, ws, MCP SDK live only in the plugin's own node_modules), so
the ratchet flagged it on 3 consecutive runs — deterministic, not a flake.
Follow the existing convention (17 plugin tests already baselined, e.g.
plugins/ruflo-adr). The real runner is `npm test` in the plugin (9/9 pass).

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* fix(x-gateway): v0.3.1 — security hardening + relay-connection optimization

Security review findings (live battery all green; static gaps fixed):
- publish(): bound msgType ([A-Za-z0-9_-]{1,64}) and payload (<=32KB) before signing
- rate limiter: evict idle buckets (10m) and cap the map (10k) — unbounded IP churn
  could grow memory without bound
- ws proxy: maxPayload 256KB on both legs; bad path now answers HTTP 404 instead
  of a bare socket destroy (Cloud Run surfaced that as 503)
Optimization:
- fetchManyOn(): several REQs over ONE authenticated connection — Seraphina now
  does 1 NIP-42 handshake per call instead of 3
- 5s TTL cache on the roster/claims resources to absorb read bursts
+1 test (10/10): bounds, bounded buckets, proxy limits, single-handshake multi-REQ.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67

* feat(x-gateway): v0.3.2 — canonical relay wss://relay.ruv.io + legacyRelay metadata

relay.ruv.io is a Cloud Run domain mapping for buzz-relay (Cloudflare CNAME, unproxied so
WebSocket upgrades go straight to Cloud Run). The old run.app host remains routable and is
advertised as legacyRelay in / and ruv://federation/registry so pinned clients keep working.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
2026-09-10 02:34:56 +00:00
ruv 13cbd697db fix(release): enforce CI and idle assignment safety 2026-08-15 11:14:06 -04:00
ruv 21f7c0adb8 fix(release): align bundled runtime metadata 2026-08-14 19:21:33 -04:00
ruv 2ec82b0cd1 chore(release): prepare stable 3.38.10 train 2026-08-14 17:35:47 -04:00
rUv 1e24bb8878 fix(agent): propagate explicit provider/model config into agent execution (#2962) (#3007)
* fix(agent): propagate explicit provider/model config into agent execution (#2962)

`providers configure` and `agent spawn --provider/--model` persisted user
intent but the execution path (callAnthropicMessages, executeAgentTask,
determineAgentModel) only ever consulted env vars and a 5-alias model list,
silently discarding both. A local Ollama/OpenRouter setup with explicit
config would either fail closed or fall back to whatever provider the
env vars happened to select.

- determineAgentModel(): treat any non-alias config.model string as an
  explicit selection via the existing modelId fast-path, instead of
  falling through to task-based routing / agent-type defaults.
- agent_spawn: forward config.provider to the stored (and returned) agent
  record when it's an unambiguous explicit choice ('ollama'/'openrouter';
  'anthropic' is excluded since the CLI silently defaults to it when
  --provider isn't passed).
- callAnthropicMessages(): accept an optional provider param and consult
  the persisted `agents.providers` config for baseUrl/apiKey/model when
  env vars are absent. A self-hosted Ollama baseUrl no longer requires
  the undocumented OLLAMA_API_KEY='local' sentinel.
- executeAgentTask(): forward agent.provider into the first dispatch call.

Precedence: explicit per-agent flag > env vars > persisted config >
key-presence inference (unchanged, last resort).

Co-Authored-By: RuFlo <ruv@ruv.net>

* fix(ci): update #2042 smoke to match the widened OpenRouter branch shape

The #2042 regression smoke statically matched the literal token sequence
`useOpenRouter && openrouterKey`. #2962 widened that condition to
`if (useOpenRouter) { const apiKey = openrouterKey || persistedOpenRouter?.apiKey; ... }`
so a persisted `agents.providers` config can supply the key when no env
var is set — the smoke's ordering guarantee (OpenRouter branch reachable
before the Anthropic-key early-return) still holds, only the literal
pattern needed updating.

Co-Authored-By: RuFlo <ruv@ruv.net>
2026-08-12 18:28:25 -04:00
rUv 83e536396f fix(scaffold): add dead-reference guard + deterministic CLI remap (ADR-382 Part C) (#2974)
Adds scripts/smoke-init-scaffold-references.mjs (ADR-382 Part C, #2971):
four static assertions over v3/@claude-flow/cli/.claude/** and the
plugin/marketplace surface, deriving the live MCP tool set and canonical
CLI form from source rather than hand-maintaining them.

  1. dead `npx claude-flow` (bare) invocation
  2. dead `mcp__claude-flow__<tool>` references not in the live registry
  3. plugin .mcp.json launches with no local-bin-first resolver (Part A regression guard)
  4. plugins/* directories missing from .claude-plugin/marketplace.json

Ships warn-only (no --strict) with a documented --strict flag to flip once
backlogs clear. Wired into .github/workflows/v3-ci.yml as a new job,
gated on plugins/*/.mcp.json, .claude-plugin/marketplace.json, and the
script itself (v3/@claude-flow/cli/.claude/** was already a trigger path).

Deterministic remap applied: 701 occurrences of bare `npx claude-flow`
across 142 files -> `npx @claude-flow/cli@latest` (mechanical, 1:1, verified
by rerunning the guard: check 1 701 -> 0, check 2 unchanged at 410). The
386 legacy `npx claude-flow@alpha` / `@v3alpha` occurrences elsewhere in
the same tree were left untouched (still-maintained dist-tags).

NOT remapped in this PR (tracked follow-up, per ADR-382 Part C's own
guidance not to guess): 410 dead `mcp__claude-flow__<tool>` occurrences
across 98 files, 54 distinct dead tool names (memory_usage: 53 files,
task_orchestrate: 41, sparc_mode: 34, swarm_monitor: 16, plus 50 more at
lower counts). Each requires per-call-site judgment about the surrounding
example's intent (store vs retrieve vs list, 1-line vs restructured
2-line orchestration calls) that a bulk regex pass cannot make safely.

Checks 3-4 are expected-red until ADR-382 Part A merges (plugin/ruflo-core
.mcp.json resolver + the 3 missing marketplace entries:
ruflo-agntcy, ruflo-bbs-federation, ruflo-business-pods).
2026-08-11 18:39:17 -04:00
rUv f35c545fbe feat(metaharness): pull in @metaharness/turn-credit + fix stale router/darwin pins (#2958)
* feat(metaharness): pull in @metaharness/turn-credit + fix stale router/darwin pins (post metaharness#176)

metaharness#176 shipped @metaharness/turn-credit (ADR-248, recursive
turn-level credit assignment) and bumped darwin to 0.9.0 (ADR-249 signal
seams) + router to 0.4.0 (calibration module). This brings ruflo's
dependency contract back in sync and makes the new package available:

- Add @metaharness/turn-credit ~0.1.0 as a new optionalDependency,
  following the same "must be installable, not peer-only" pattern
  darwin/flywheel/radio already use (ADR-150) — dependency-free, 64.9K
  unpacked, zero lifecycle scripts, same profile as the other three.
- Bump @metaharness/darwin ~0.8.3 -> ~0.9.0 and the paired MH_DARWIN_PIN
  constant in distill-oracle.ts (was already out of range: tilde only
  absorbs patches, and 0.9.0 is a minor bump).
- Bump @metaharness/router peer range ^0.3.2 -> ^0.4.0 (still deliberately
  peer-only + triple-gated behind CLAUDE_FLOW_ROUTER_NEURAL=1, per
  neural-router.ts — that design choice is unchanged, only the stale
  range is fixed) and update the matching manual-install hint in neural.ts.
- scripts/check-metaharness-pins.mjs + scripts/metaharness-clean-install-test.mjs:
  add turn-credit to the watched/contract-checked package list.
- .github/workflows/no-cli-optdep-bloat-2561.yml: CLI_MAX 10 -> 13. The
  prior bump (PR #2956) left zero slack (budget == count exactly), which a
  code review flagged as a latent trap — the very next unrelated optional
  dep would trip this guard. This bump leaves real headroom (11 declared
  today, budget 13) instead of repeating that mistake.
- Also fixes a live ReferenceError in distill-oracle.test.ts (MH_DARWIN_PIN
  used but never imported) — a gap from the prior release's test fix that
  somehow didn't surface in that PR's CI; caught here while touching the
  same file.

Co-Authored-By: RuFlo <ruv@ruv.net>

* fix: regenerate v3/pnpm-lock.yaml — was out of sync with package.json edits

The previous commit edited v3/@claude-flow/cli/package.json directly
(darwin/router/turn-credit pin changes) without regenerating the pnpm
workspace lockfile, so every CI job running `pnpm install --frozen-lockfile`
failed immediately with ERR_PNPM_OUTDATED_LOCKFILE — cascading into every
downstream smoke/test job that depends on that install step.

Regenerated with pnpm@8.15.9 (matching CI's pinned version) in an isolated
worktree to avoid the lockfile-version drift a newer local pnpm would
introduce.

Co-Authored-By: RuFlo <ruv@ruv.net>
2026-08-10 15:08:06 -02:30
rUvandClaude 0b3cfb77d6 MetaHarness hardening: repair the dependency contract + strict sequential promotion evidence (#2956)
* feat(metaharness): repair dependency contract + strict sequential promotion evidence

Item 1 — dependency & compatibility repair (the contract was silently broken):
- @metaharness/darwin ^0.8.3, @metaharness/flywheel ^0.1.10, and
  @metaharness/radio ^0.1.0 are now explicit optionalDependencies of
  @claude-flow/cli (optional PEER deps are never auto-installed, so a clean
  ruflo install shipped with zero MetaHarness packages on disk)
- check-metaharness-pins.mjs now searches dependencies, optionalDependencies,
  AND peerDependencies; UNDECLARED and PEER-ONLY (for installable pins) are
  fatal drift instead of silently passing; radio added to the watch list;
  --require-installed asserts the real published symbol contracts
  (RefineMutator, withSequentialEvidence, RadioBus, ...)
- new scripts/metaharness-clean-install-test.mjs: installs the declared
  ranges into a pristine temp dir and asserts every advertised export;
  wired as a MANDATORY clean-install job in metaharness-ci.yml
- doctor: new "MetaHarness declared packages" check FAILS (not warns) when a
  declared optional dep does not resolve at runtime; --component metaharness
  now runs upstream + declared-deps + integration checks
- MH_DARWIN_PIN bumped 0.8.0 → 0.8.3 in lock-step with the declared range

Item 2 — strict promotion evidence at the transaction authority:
- receipts now carry task-level pairedOutcomes (taskId + per-task baseline/
  candidate scores) behind heldOutDeltas; verifyFlywheelReceipt refuses rows
  that cannot reproduce their aggregate; evaluateFlywheelCandidate populates
  them; pre-existing receipts still verify byte-identically
- new flywheel-sequential-evidence.ts: anytime-valid e-process over
  discordant paired outcomes (testing-by-betting) composed with per-candidate
  alpha allocation (alpha_k = alpha_total * 6/(pi^2 k^2)), so the family-wise
  false-promotion probability across an ADAPTIVE candidate stream is bounded
  by alpha_total = 5%
- promoteFlywheelCandidate requires paired evidence by DEFAULT — aggregate-
  only receipts are refused, never silently downgraded (the upstream
  withSequentialEvidence fallback hole); explicit
  --allow-aggregate-evidence / allowAggregateEvidence escape hatch for
  pre-upgrade receipts; alpha spend is persisted per receiptId in the
  transaction state (looking spends alpha; retries reuse their index)
- acceptance test: 1,000 null-improvement streams (20 adaptive candidates x
  40 worst-case discordant pairs) — measured family-wise false promotion
  0.6%, within the <=5% budget

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy

* chore: gitignore node-compile-cache build artifact

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy

* fix(ci): satisfy #2561 guard, dep-overlap audit, and smoke contract for metaharness optional deps

Three CI failures from the dependency-contract repair, three fixes:

- tilde-pin @metaharness/{darwin,flywheel,radio} (~0.8.3 / ~0.1.10 / ~0.1.0)
  per the ADR-150 anti-caret rule enforced by smoke step 17z73 — upstream is
  unstable, tilde absorbs patches but never minors
- remove the three from peerDependencies/peerDependenciesMeta: an entry in
  BOTH optionalDependencies and peerDependencies crashes npm 11.x arborist
  on dedupe (#1147/#2018 dep-overlap audit); optionalDependencies alone is
  the declaration that actually installs
- update the #2561 cold-startup guard per its own escape clause: budget
  8 → 10 and drop @metaharness/darwin from the forbidden list — that entry
  dated from the pre-0.8 heavy-tree era; darwin@0.8.3 is 1.8M unpacked with
  ZERO dependencies and no lifecycle scripts (all three packages combined:
  2.3M, 753ms cold install into an empty dir, measured 2026-08-10; the
  clean-install CI job re-verifies on every pin change)
- smoke.sh 17h now accepts the componentMap ARRAY form
  ('metaharness': [checkMetaharness, ...]) introduced with the
  declared-packages doctor check

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy

* fix(ci): remove apostrophe that broke the single-quoted guard script

The #2561 guard runs as `node -e '...'` inside bash; an apostrophe in a
comment ("guard's") terminated the quoted string and bash tried to execute
the next // comment line (exit 126, '//: Is a directory'). Verified by
executing the extracted run block end-to-end: all three checks OK, exit 0.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy

* feat(flywheel): ADR-381 — alpha-stream governance, reset epochs, accept/v2+seq

Completes the sequential-evidence governance PR #2956 opened:

ADR-381 (new) records the decisions:
- the ADR-322 promotion ledger IS the alpha stream (per project root) —
  receipt lineageIds default to fresh UUIDs and cannot scope the control
- evidence epochs: resetSequentialEvidence({confirm, reason}) archives the
  spend into an append-only sequentialResets audit trail, EXPIRES every
  outstanding evaluated receipt (fresh-data enforcement — old-epoch
  evidence cannot be replayed against a reopened budget), and increments
  evidenceEpoch; surfaced as `flywheel evidence-reset --reason … --confirm`
  (CLI) and the metaharness_flywheel MCP op, both behind the same policy
  gate as promotion
- accept/v2+seq: the ADR-176 generations loop now decides each bundle with
  a third conjunct — the e-process over the bundle's own embedded per-task
  holdout at alpha_k for its position in the attempts stream — with the
  full verdict recorded in the bundle and independently replayed by
  verifyReceiptBundle (v1 bundles keep v1 semantics; versions pin per
  bundle); flywheelStatus surfaces next test index / threshold / minimum
  pairs / remaining budget so exhaustion reads as plateau, not mystery
- two-layer pre-flight: evaluateFlywheelCandidate annotates (never blocks)
  when the promotion holdout cannot clear the next threshold on a perfect
  sweep; promoteFlywheelCandidate refuses size-inviable receipts BEFORE
  allocating an alpha index — sample size is ancillary, so the refusal
  looks at no evidence and spends no budget

Tests: reset semantics (archive/expire/epoch/fresh-promote), alpha-free
size refusal vs alpha-spending e-process refusal, v2 promote/reject/replay
including tamper + index-shopping detection, v1 compat, helper math
anchors (min pairs = 9 at test 1), budget monotonicity. 82 tests green
across the eight affected suites; tsc build clean.

Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy

* fix(flywheel): close 3 concurrency/epoch gaps in ADR-381 sequential evidence

Code review of this PR found three real correctness bugs that can silently
violate the <=5% family-wise false-promotion guarantee that is the PR's own
core deliverable:

- flywheel-transaction.ts: resetSequentialEvidence's "fresh data only"
  guarantee only expired receipts already registered at reset time; a
  receipt whose evidence predates a reset but is registered afterward was
  silently admitted into the new, cheaper epoch (index shopping). Now
  tracks evidenceEpochStartedAt and promoteFlywheelCandidate refuses any
  receipt whose payload.issuedAt predates it, before allocating an alpha
  index.

- harness-flywheel-generations.ts: the daemon generations loop computed its
  sequential-evidence testIndex via an unlocked loadAttempts(root).length+1
  read before appending. Two overlapping runFlywheelGeneration calls on the
  same root could be assigned the same test index and spend the same
  alpha_k twice. The read-index -> build-bundle -> append critical section
  now runs under the same O_EXCL lock pattern flywheel-transaction.ts
  already uses.

- harness-flywheel.ts: evaluateFlywheelCandidate's sequentialPreflight
  (backing the `promotable` flag) was computed from an unlocked snapshot
  taken before the async retrieval/scoring work, so it could go stale by
  the time promoteFlywheelCandidate allocates the real index under its own
  lock. Moved to the latest possible read (after receipt registration) to
  minimize the window; documented as advisory, since promoteFlywheelCandidate
  remains the sole authority.

Added a regression test for the epoch-boundary refusal and a concurrency
test proving two racing generations get distinct sequential test indices.

Co-Authored-By: RuFlo <ruv@ruv.net>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 11:32:40 -02:30
rUv fabcc9261a fix(hooks): Codex hooks.json schema + PreToolUse verdict compat (#2857)
Codex's plugin hook-manifest loader accepts only `description` and
`hooks` at the top level and hard-rejects the rest of the manifest on
anything else. plugins/ruflo-core/hooks/hooks.json and
plugins/ruflo-cost-tracker/hooks/hooks.json carried `_note` /
`_platform_note` documentation fields, so a fresh Codex install of
ruflo-core@ruflo (still true as of the currently-published 3.32.38 /
ruflo-core 0.2.5) fails to load the plugin at all with:

  failed to parse plugin hooks config .../hooks.json:
  unknown field `_note`, expected `description` or `hooks`

Fold the doc content into `description` (content preserved, no hook
commands changed) for both marketplace-listed manifests, and add a
strict top-level-key check to
scripts/audit-plugin-hooks-cross-platform.mjs so this class of
regression fails CI going forward. `.claude-plugin/hooks/hooks.json`
is POSIX-only and not marketplace-distributed (not what Codex fetches
for ruflo-core@ruflo); its doc fields are folded too but its
audit-script flags (`_platform`, `_legacy_unaudited_shim`) are kept.

Second, deeper bug once the manifest loads at all: `modify-bash`/
`modify-file` PreToolUse hooks always echo Cursor's
`{"permission":"allow"}` verdict. Codex's own strict output schema
(additionalProperties: false) rejects that shape outright and reports
"hook returned invalid pre-tool-use JSON output" on every single tool
call — verified directly against the real parser
(codex-rs/hooks/src/engine/output_parser.rs,
codex-rs/hooks/src/events/pre_tool_use.rs) and its generated JSON
schema. Empty stdout is the one input Codex's parser treats as
no-opinion/implicit-allow with no error. `isCodexPluginHost()` (added
for #2816, ruflo-core 0.2.5) already detects Codex via its
PLUGIN_ROOT/PLUGIN_DATA env vars and correctly suppresses the verdict
in that case — harden it with a `turn_id`-based fallback (a documented
Codex-only field always present in Codex's PreToolUse input JSON, per
codex-rs/hooks/schema/generated/pre-tool-use.command.input.schema.json)
so detection doesn't depend solely on those env vars being set.

Bump ruflo-core 0.2.5 -> 0.2.6 and ruflo-cost-tracker 0.26.2 -> 0.26.3
so Codex's per-version plugin cache invalidates and re-fetches the
fixed manifest instead of serving a stale cached copy indefinitely.

Fixes #2855, #2856.
2026-07-29 21:36:20 -04:00
rUv 401e02d511 fix: complete reports and consistent initialization for v3.32.37 (#2851)
* fix(metaharness): preserve readiness verdict payloads

* test(metaharness): cover blocked genome verdicts

* fix(adr): parse bullet metadata and relationships (#2659)

* fix(adr): align adr-create with AgentDB schema (#2651)

* fix(adr): make index updates idempotent (#2660)

* fix(memory): bound session-end graph consolidation (#2628)

* fix(memory): align active row visibility (#2652)

* fix(memory): honor database path during init

* fix(hooks): keep all shim fallback tags aligned

* fix(codex): omit unbacked full-template skills

* fix(init): generate complete native dual projects

* test(memory): isolate path and legacy-row regressions

* chore(release): prepare v3.32.37
2026-07-29 15:40:17 -04:00
rUv 67ff9898c0 fix: resolve current runtime and verification defects for v3.32.36 (#2850)
* fix(cli): make runtime status and learning signals truthful

* fix(metaharness): accept padded scan severities

* fix(codex): make core hooks and status skill native

* fix(security): make witness verification hermetic

* fix(metrics): expose grounded learning outcomes

* fix(hooks): deduplicate project and plugin events

* fix(verification): preserve security audit payloads

* fix(runtime): stabilize dual memory skills and MCP schemas

* fix(cli): harden helper signing and reflexion health

* test(release): align helper and container invariants

* chore(release): prepare v3.32.36

* test(funnel): align gates with cold-start seed pool

* fix(ci): document MCP environment precedence
2026-07-29 15:08:51 -04:00
rUv 9cf769ccce feat: ship adaptive swarm and resolve top runtime issues (#2848)
* feat: ship adaptive swarm and issue fixes

* fix: register intentional runtime escape hatches
2026-07-29 13:25:22 -04:00
rUv ddd27576cb fix(release): ship Capability Brain as a self-contained 3.32.30 train (#2829)
* fix(release): bundle policy and Codex runtimes

* fix(release): verify self-contained three-package archives
2026-07-29 01:30:53 -04:00
rUv b08246b04a chore(release): bump to 3.32.27 (#2823)
* chore(release): bump to 3.32.27

* chore(release): publish policy security dependency

* chore(release): sign 3.32.27 helper manifest

* ci: harden policy release gates
2026-07-28 22:38:57 -04:00
rUv 0e3d412756 feat(flywheel): implement ADR-322 promotion loop (#2817)
Adds verified evaluation receipts, atomic compare-and-swap promotion, bounded Darwin/local proposer integration, CLI/MCP surfaces, ADR specifications, and the 3.32.26 release bump.
2026-07-28 15:23:57 -04:00
rUv db76d67235 fix(metaharness): pin @metaharness/darwin + pin-drift guard (ruflo analog of upstream #142/#149) (#2813)
* fix(metaharness): pin @metaharness/darwin + add pin-drift guard (ruflo analog of #142/#149)

Upstream agent-harness-generator #142 (darwin caret-locked to 0.2.x, 3 majors
behind) and #149 (META_PROXY_VERSION pinned with no watcher) are both silent
pin-drift failures. Ruflo had the same latent gap on its own metaharness deps:

- distill-oracle.ts invoked `npx --yes @metaharness/darwin` with NO version pin,
  so the Tier-1 mechanical oracle floated to whatever npm `latest` was — a
  breaking darwin release could change eval behavior mid-run. Pinned via a new
  `MH_DARWIN_PIN = '0.8.0'` constant used by all three npx call sites.
- Declared `@metaharness/darwin: ^0.8.0` in optionalDependencies (subprocess-
  invoked, kept optional per ADR-150/321) so the pin has a single source of truth
  and installs cache-warm.
- Fixed a stale `@metaharness/darwin@~0.3.1` doc comment (darwin is 0.8.0).

New guard (the #149 analog ruflo lacked):
- scripts/check-metaharness-pins.mjs — diffs each declared range (metaharness,
  @metaharness/router, @metaharness/darwin) against npm `latest`, plus a
  lock-step check that MH_DARWIN_PIN satisfies the declared darwin range. Exit 1
  on drift; network flakes surface as a warning, never a false positive.
- .github/workflows/metaharness-pin-drift.yml — runs the guard weekly + on PRs
  touching the pins; opens/updates a tracking issue when a pin falls behind,
  and hard-fails PRs that introduce drift.

All pins are current (guard exits 0); router API-surface compat 9/9; CLI builds
clean. Lockfile reconciled for the new optionalDep (frozen-lockfile verified).

Refs: ruvnet/agent-harness-generator#142, #149; ADR-150, ADR-321.

Co-Authored-By: RuFlo <ruv@ruv.net>

* chore(release): bump to 3.32.25 (metaharness pin-drift guard)

Co-Authored-By: RuFlo <ruv@ruv.net>
2026-07-28 11:10:05 -04:00
rUv 27410d402b feat(security): ADR-320 — MCP Composition Inspector v2 + ChannelGuard v2 (#2791)
Follow-up from dream-cycle issue #2783 / PR #2784. Implements the v2 that the already-shipped v1 (commits 381b7ebcc/581cd2bf3) explicitly deferred: SimHash-based cross-tool fragment detection for MCP composition, and a ChannelGuard reusing the real InputValidator instead of a reinvented catalog. 29 new tests, 134/134 existing hooks tests passing, 0 regressions, clippy clean.
2026-07-27 17:03:26 -07:00
ruv 469a901eda fix(bridge): wire bridgeRecordFeedback to real intelligence.recordTrajectory (#2786 fix-3)
The prior code called `learningSystem.recordFeedback`/`.record` and
`reasoningBank.recordOutcome`/`.record` — none of those methods exist
on the LocalSonaCoordinator / LocalReasoningBank instances the bridge
actually wires into the registry (see initializeIntelligence in
memory/intelligence.ts). Two silent `catch { /* API mismatch — skip */ }`
blocks swallowed every call. Feedback recording was 100% a no-op for the
CLI's default in-process intelligence path.

Real fix: call `intelligence.recordTrajectory(steps, verdict)` — the
same public API `hooks_post-command` already uses. It initializes
lazily, embeds the step, drives SONA + pattern distillation.

Also fixed the ReasoningBank pattern-store branch to call the ONE
method that actually exists on LocalReasoningBank (`.store(pattern)`)
with the correct StoredPattern shape.

E2E verified in a fresh scratch cwd (v3.32.10 CLI):
  - `memory init`
  - 3x `hooks post-task --task ... --store-results true`
  - `hooks intelligence stats` — Neural Persistence reports 6 trajectories
    on disk + 3 pattern entries. Before this fix: 0.

No mocks, no silent catches on the happy path.

Closes: #2786 fix-3 (the last flagged item from the 2026-07-26 tracker
sweep). Was tagged "architectural" because the sweep agent assumed the
target was the AgentDB LearningSystem/ReasoningBank; turned out the
registry wires the local intelligence classes instead, and the intelligence
module already exposes the correct public API. One-file surgical fix.

Co-Authored-By: RuFlo <ruv@ruv.net>
2026-07-26 18:47:59 -04:00
ruvnet 9810d8d9c1 fix(statusline): real model name from stdin + worktree version resolution
Fixes #2733, #2742.

#2733 — hooks.ts's getUserInfo() hardcoded
`const modelName = 'Opus 4.6 (1M context)'`, ignoring the actual active
model Claude Code passes on stdin entirely. Cosmetically masked in the
default render path (the generated .claude/helpers/statusline.cjs
already parses stdin correctly and overrides this), but real for
direct/manual `hooks statusline` CLI use or any stdin-parse failure in
the wrapper. Fixed by mirroring statusline.cjs's own getModelFromStdin()
approach inside hooks.ts, so the CLI subcommand is correct standalone.

#2742 — getPkgVersion() in the generated statusline.cjs only probed
CWD-relative paths (CWD/node_modules/..., CWD/v3/@claude-flow/cli/...).
A linked git worktree has no node_modules of its own (worktrees don't
get their own `npm install`), so every probe missed and the version
silently fell back to the baked-in default from whenever the helper was
last generated. Fixed with a pure-fs worktree-root resolver: a linked
worktree's `.git` is a plain FILE containing `gitdir: <main>/.git/
worktrees/<name>`; walk up from CWD, parse the pointer, strip the
trailing segment to recover the main repo root, and probe its
node_modules/v3 paths too. No `git rev-parse` spawn (statusline renders
are latency-sensitive). Caught and fixed a real bug in this same
resolver during testing: git writes the gitdir pointer with forward
slashes even on native Windows, so a path.sep-based (backslash) marker
search silently never matched — normalize to forward slashes before
searching.

`v3/@claude-flow/cli/.claude/helpers/statusline.cjs` is the source of
truth per the #2679 redesign (generateStatuslineScript() reads it and
substitutes two tokens); regenerated the propagated root-level copy via
scripts/regen-statusline-artifact.mjs, which also needed a small
cross-platform fix (dynamic import() of a raw Windows path isn't a
valid ESM specifier — wrap with pathToFileURL()).

Also corrected a stale/misleading comment in message-transport.ts
claiming a fresh install "shows the in-code fallback pool until the
first refresh lands" — there has never been an in-code fallback pool
since ADR-311 ("zero local promo content"); a fresh install's promo
row is genuinely empty until the first background refresh lands.
Found while investigating a "promo doesn't show on new installs"
report; the underlying fail-closed design and SessionStart-triggered
refresh mechanism were confirmed working as intended (this machine's
own ~/.ruflo state shows a healthy, recently-rotated promo history) —
only the comment was wrong, not the behavior.

Left alone (separate, unreferenced, ~500-line legacy implementation
predating the #2195 delegation rewrite, no source references in its
own package): v3/@claude-flow/mcp/.claude/helpers/statusline.cjs.

Verified: 21/21 new + existing statusline tests pass, 122/122 funnel
tests pass, 10/10 hooks tests pass, clean tsc --noEmit. Manually
confirmed #2733 (real stdin model name renders; malformed/empty stdin
falls back to "Claude Code", never the old hardcoded string) and #2742
(a real `git worktree add` scenario resolves the main repo's version
instead of falling back) end-to-end.
2026-07-20 15:58:21 -04:00
ruvnet 0a110aee9f fix(audit): register #2721's test-only hook env vars as escape hatches
RUFLO_HOOK_CLI_OVERRIDE and RUFLO_HOOK_DEBUG_STDOUT (both added to
plugins/ruflo-core/scripts/ruflo-hook.cjs to let test-hooks.mjs point
at a local CLI build and observe its output) tripped the
env-var-precedence audit's "CLI flag must win" requirement. Same
category as the existing RUFLO_HOOK_SKIP_NPX entry: hook scripts have
no CLI-flag surface to attach to (invoked by hooks.json, never a
user-typed command), and both are test-only — production never sets
them.
2026-07-18 20:04:50 -04:00
ruvnet b68ad4ccba fix(plugins): make ruflo-core/ruflo-cost-tracker hooks Windows-native (#2721)
Both plugins' hooks.json wrapped every command in `/bin/bash -c '...'`,
which fails outright on native Windows (no such path) -- Codex/Claude
Code report "PreToolUse hook (failed) -- exit code 1" on every tool
call. The `_platform: posix` / "ruflo init overrides this on Windows"
claim in both files was never actually true: Claude Code merges
plugin-declared hooks additively with any init-generated
.claude/settings.json, it doesn't replace them, and there's no `ruflo
init` step at all in the reported Codex marketplace install flow.

Fix: every hook command is now a `node -e` bootstrap that resolves
plugins/*/scripts/ruflo-hook.cjs from process.env.CLAUDE_PLUGIN_ROOT
inside Node -- no shell env-var expansion (${VAR} vs %VAR%), so the
exact same command string runs unchanged on Windows/macOS/Linux.

ruflo-core's ruflo-hook.cjs (previously a full port of ruflo-hook.sh
that existed on disk but was never referenced by hooks.json) gained:
  - JSON parsing of the hook event from stdin (replaces jq) for
    post-command/post-edit, deriving the same CLI flags the bash
    version computed
  - the PreToolUse permission-allow stdout echo Cursor's stricter
    contract requires (previously only the bash wrapper's trailing
    printf did this)
  - precompact-manual/precompact-auto guidance text (previously plain
    bash echoes, no CLI call)
  - a real Windows shell-quoting fix: shell:true with an args array
    does NOT quote array elements, so "echo hi" silently truncated to
    "echo" and a heredoc's `<<` errored as unexpected -- skip the
    shell entirely for `node` invocations (never a .cmd shim, so
    CreateProcess gets the argv array byte-for-byte)

cost-tracker's existing ruflo-hook.cjs (already correct, just
orphaned) needed no logic changes, only wiring.

Also:
  - corrected the false "_platform_note" claims about ruflo init
    overriding plugin hooks
  - hardened scripts/audit-plugin-hooks-cross-platform.mjs: a
    POSIX-exempt hooks.json now must actually reference its sibling
    .cjs shim, not just have one sitting on disk unreferenced (which
    is exactly the shape cost-tracker shipped in undetected)
  - added windows-latest to the plugin-hooks-smoke CI matrix (it was
    ubuntu/macos-only because the old bash-based hooks.json couldn't
    run on Windows at all) and rewrote test-hooks.mjs to drive hooks.json's
    literal command strings via `shell: true` -- exactly how Claude
    Code/Codex invoke them -- instead of wrapping everything in an
    explicit `bash -c` that could never have caught this bug
  - flagged (not fixed) a separate, currently-published, actively
    maintained plugin package (.claude-plugin/ + plugin/, the older
    "claude-flow" plugin, not listed in the ruflo marketplace) with
    the same underlying bug via jq/xargs pipes instead of bash --
    explicitly marked _legacy_unaudited_shim so the hardened audit
    doesn't silently regress on out-of-scope work

Verified locally on native Windows (this fix's actual target
platform): all 17 ruflo-core hook cases pass, all 3 cost-tracker
cases pass, the existing 12-case smoke-ruflo-hook-cjs.mjs passes
unchanged, both hook-command audits pass clean.

Fixes #2721
2026-07-18 19:05:06 -04:00
ruvnet a4a7d99c22 test(statusline): document hook-only identity setting 2026-07-16 23:31:05 -04:00
ruvnet 1fb874005c test(plugins): preserve standalone MCP catalog assertions 2026-07-16 23:20:35 -04:00
ruvnet 5e66f065e9 test(plugins): align namespace and stable hook shims 2026-07-16 23:14:09 -04:00