mirror of
https://github.com/ruvnet/ruflo.git
synced 2026-09-28 14:32:58 +08:00
main
232
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8761c96a27 |
fix(ci): warm up the AgentDB registry before ADR-130 P3's timed assertions
graph-trajectory-hooks-smoke failed 3/3 times in CI on unmodified #3467 code (reverted my own unrelated attempt at a fix in the process — see prior commit). Root cause identified: bridgeRecordFeedback()'s first-ever call opens the AgentDB ControllerRegistry (new ControllerRegistry(). initialize({ embeddingModel: ... })), which resolves/attempts the ONNX embedding pipeline before falling back to mock embeddings — a real, one-time-per-process bootstrap cost. TEST 1 (trajectory-step on a never-started trajectory) never reaches that code path, so TEST 2 (post-task) was the first and only call to pay it, entirely inside its own <200ms budget: ~1000-1250ms on GitHub Actions' cold filesystem cache, 20-46ms locally with a warm one. Added an explicit, untimed warm-up call before TEST 1 begins so the one-time cost is paid during setup, matching how a real long-lived Ruflo session actually behaves (bootstrap once at startup, not per call). Verified 3/3 clean local runs post-fix, TEST 1 now measuring 1-2ms (down from paying the same hidden cost intermittently). |
||
|
|
0d0ebbe305 |
fix(ci): add two new node:test-only ADR files to the vitest ratchet baseline
#3432 and #3437 add plugins/ruflo-adr/scripts/__tests__/import-root-exit-3097.test.mjs and verify-read-safety-3147.test.mjs, both node:test-only (unlike parser-relations-3096.test.mjs's dual-mode shim). Root vitest can't discover a test suite in either, which the CI ratchet correctly flagged as 2 new unexpected failing files. Same documented class as the plugin's four other baselined test files (see the file's own comment on plugins/ruflo-x-gateway). |
||
|
|
24a80bd75d | test(ci): isolate witness fixture and search mocks | ||
|
|
9f3ac3b801 | Merge branch 'pr-3412' into integ/train-3.45.1 | ||
|
|
4a6e063de7 | Merge branch 'pr-3391' into integ/train-3.45.1 | ||
|
|
72a39572f2 | Merge branch 'pr-3383' into integ/train-3.45.1 | ||
|
|
2f192d4eb9 | Merge branch 'pr-3367' into integ/train-3.45.1 | ||
|
|
227af0b969 | Merge branch 'pr-3424' into integ/train-3.45.1 | ||
|
|
4c2c62df27 |
dream(intelligence, evaluated): wire CLAUDE_FLOW_PRIOR_DECAY env override for ModelRouter's dormant decay primitive (#3349) (#3350)
* dream(intelligence): #3349 wire CLAUDE_FLOW_PRIOR_DECAY env override for ModelRouter's dormant decay primitive ModelRouter's discounted-Thompson-sampling priorDecay (built/tested/ benchmarked in #3049) shipped permanently inert: DEFAULT_CONFIG.priorDecay was hardcoded to 1 (disabled) with no env/config knob, unlike its sibling maxUncertainty (envMaxUncertainty()). Adds envPriorDecay() mirroring that exact pattern; default behavior is unchanged when the env var is unset. Evaluation: stash-isolated discriminating test (1/3 new tests fails against reverted source, 16/16 pass restored); full @claude-flow/cli suite and tsc --noEmit both byte-identical baseline vs candidate outside the 3 new tests; benchmark re-run reproduces the 2026-08-17 receipt exactly. Real non-stationary recovery win in the 'low' bucket; a small but statistically real stationary-accuracy cost in the 'med' bucket, accepted only under the pre-existing ±1pp tolerance band from the original receipt — disclosed in the issue/gist rather than oversold. Independent adversarial critic verdict: CONFIRMED-WITH-CAVEATS. Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_011tXcJWNao4vh57f4uFXutg * docs(dream-cycle): append 2026-09-17 ledger row + backfill 09-13..09-16 gap/run verification Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_011tXcJWNao4vh57f4uFXutg * fix(audit): register CLAUDE_FLOW_PRIOR_DECAY as a documented env-var escape hatch scripts/audit-env-var-precedence.mjs (ADR-125/ADR-130) flagged the new CLAUDE_FLOW_PRIOR_DECAY read in model-router.ts as undocumented CLI-flag precedence. It's the same operator-knob shape as the adjacent, already- registered CLAUDE_FLOW_MAX_UNCERTAINTY (envMaxUncertainty() is the pattern envPriorDecay() mirrors) — no CLI invocation owns the router's persisted lifetime, so there's no flag to wire precedence against. Registered alongside it with the same justification. Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_011tXcJWNao4vh57f4uFXutg --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
0702af8158 | fix(witness): fail marker-preserving hash drift in strict mode | ||
|
|
abf6da6beb | test(security): distinguish direct argv from shell splitting | ||
|
|
2950997f3e | test(security): exercise every shipped github-safe helper | ||
|
|
82dac70fd7 |
feat(router): ADR-389–391 — refresh router helper, opt-in MiniLM embedder, benchmark gate (no default change) (#3410)
* docs(adr): ADR-389..391 — router: refresh the helper, use the real embedder, change defaults only on a measured win ADR-389: add router.js to CRITICAL_HELPERS + re-sign, so installed copies get the word-boundary fix (#3401); today they are frozen at first write. ADR-390: hooks_route embeds patterns and tasks with the local MiniLM chain (generateLocalEmbedding), hash embedder as the reported fallback. ADR-391: frozen, labelled 150-200 prompt corpus; A/B/C/D (current, MiniLM, typesafe-hash, typesafe-onnx); ADR-150-style AND-gate for promotion. Co-Authored-By: RuFlo <ruv@ruv.net> * feat(init): refresh router.js as a critical helper on upgrade (ADR-389, #3401) Installed copies of .claude/helpers/router.js never received the #2257 word-boundary fix: helper-refresh only re-copies CRITICAL_HELPERS and router.js was not on the list, so an April substring router kept routing "sync and review latest issues" to tester. - helper-refresh.ts: add router.js to CRITICAL_HELPERS; generator fallback also emits it (generateAgentRouter). - sign-helpers.mjs / verify-helpers.mjs / executor.ts executeUpgrade / smoke-helper-signing-security.mjs: keep their copies of the list in sync. - New test: stale April router.js is replaced by the package copy via the signed refresh path (throwaway key), after which the prompt no longer routes to tester. helpers.manifest.json is NOT re-signed here. Until it is, verify-helpers fails ("manifest has no entry for router.js") and the production refresh is fail-closed blocked for all helpers. Co-Authored-By: RuFlo <ruv@ruv.net> * bench(router): ADR-391 labelled evaluation corpus (197 prompts, frozen) Blind-labelled corpus for comparing routing candidates: 11 labels incl. none, 64 adversarial trap cases, deterministic stratified dev/test split (i%5 in {0,2} -> dev), dependency-free validator that recomputes the split and prints the sha256. Frozen at sha256 b0c1923b2813b61907304dff56c53d199b709797bfb9600a7c1f496cf0c0bf04. Co-Authored-By: RuFlo <ruv@ruv.net> * feat(hooks): ADR-390 — semantic router can use the real MiniLM embedder hooks_route compared tasks with pattern keywords using a character hash (generateSimpleEmbedding), which measures spelling, not meaning. This adds a router-embedder abstraction (src/ruvector/router-embedder.ts): - embedForRouter(texts, kind) -> { vectors, embedder: 'minilm'|'hash', reason? } - MiniLM path uses generateLocalEmbedding ONLY (never the bridge-first generateEmbedding, #2312). backend !== 'onnx', a throw, or a non-384-d vector degrades EVERY text in the call to the hash, with a reason. - Selection: CLAUDE_FLOW_ROUTER_EMBEDDER=minilm|hash; DEFAULT_ROUTER_EMBEDDER stays 'hash' until ADR-391's benchmark decides. getSemanticRouter builds both the native VectorDb and pure-JS indexes from one embedder, records which one actually built the index, embeds the query with that same embedder, rebuilds when the requested embedder changes, and shares one in-flight build between concurrent callers. If the query fails on MiniLM after a MiniLM index was built, the index is rebuilt with the hash so the two spaces never mix. hooks_route results gain `embedder` (+ `embedderReason` when degraded); existing fields are unchanged. routeTaskForBench(task, { embedder }) is an internal, bench-only entry that runs the same local path as hooks_route (skipping the AgentDB pre-route and the opt-in typesafe wrapper) for ADR-391. Co-Authored-By: RuFlo <ruv@ruv.net> * chore(audit): register CLAUDE_FLOW_ROUTER_EMBEDDER as a known env escape hatch ADR-390's router embedder selector is process-lifetime MCP router state, like CLAUDE_FLOW_DISABLE_NATIVE_ROUTER; the ADR-391 bench passes the embedder explicitly rather than through a CLI flag. Co-Authored-By: RuFlo <ruv@ruv.net> * bench(router): ADR-391 first run — no candidate promoted; fix typesafe-router test (e) Benchmark (run-bench.mjs) on the frozen 197-prompt corpus (sha256 b0c1923b…), test split n=113: A current (hash) 25.7% acc, p95 1.06 ms B MiniLM (ADR-390) 36.3% acc, p95 6.21 ms (+10.6 pts, +485% p95) C typesafe hash 26.5% acc, p95 1.20 ms D typesafe onnx 29.2% acc, p95 8.35 ms (raw pick 44.3%, gated to 5.6%) No candidate passes the ADR-391 AND-gate (B fails the relative latency bound), so DEFAULT_ROUTER_EMBEDDER stays 'hash' and typesafe stays opt-in. Receipt committed under benchmarks/router/results/. ADR-391 gains a Results section with the findings (structural ceiling: no pattern returns researcher/reviewer/none; latency criterion mis-specified for a ~1 ms base; typesafe gate too tight) as owner decisions, not applied. ADR-389 -> Implemented; ADR-390 -> Implemented as opt-in; ADR-391 -> Implemented. typesafe-router test (e) was red on main since #3402 made the keyword fallback word-bounded: it still asserted the old "latest"->tester fallback. It now asserts the typesafe override and that a legacy fallback is kept. Co-Authored-By: RuFlo <ruv@ruv.net> |
||
|
|
72e17533ca |
fix(memory): release graph-edge-writer's native WAL handle so sql.js memory_store isn't refused (#3397) (#3405)
* fix(memory): release graph-edge-writer's native WAL handle so sql.js memory_store isn't refused (#3397) graph-edge-writer cached its better-sqlite3 WAL handle in a module singleton for the whole MCP server lifetime; nothing but the test-only _resetBridgeDb ever closed it. Its -wal/-shm sidecars therefore stayed on disk forever, and the #2735 sidecar guard (correctly) refused every later sql.js whole-image write. On Windows, where the native AgentDB bridge is off by default (#3024), that turned one hooks_post-task into a permanent memory_store outage for the server. - getBridgeDb() arms an unref'd idle timer (1s, CLAUDE_FLOW_GRAPH_EDGE_IDLE_MS) that checkpoints (TRUNCATE) and closes the handle; re-armed on every use. - New releaseBridgeDb(dbPath?) is called by memory-initializer right before each #2735 guard, so a store immediately after an edge write works and the edge is checkpointed into the image sql.js rewrites. - process 'exit' hook releases the handle so a clean shutdown does not leave stale sidecars for the next process. No signal handlers (would change Node's default termination). - The guard itself is unchanged: an open WAL connection makes a whole-image write unsafe even in-process until its WAL is checkpointed. Fixes #3397 Co-Authored-By: RuFlo <ruv@ruv.net> * chore(audit): register CLAUDE_FLOW_GRAPH_EDGE_IDLE_MS as an env-only tuning knob (ADR-125) Co-Authored-By: RuFlo <ruv@ruv.net> |
||
|
|
1116724a65 |
ci(guard): make the leaf-publish diff check actually run in PR CI
The guard's first real CI run passed but reported:
NOTE: diff check NOT EXERCISED — could not diff against 'origin/main' (shallow clone?)
So only the range check ran; check 2 — the half that replays #3335/#3390 by
catching a leaf source change with no version bump — sat out entirely. The
NOT EXERCISED note is what made that visible instead of reading as a pass,
which is precisely why a skipped check must never print like a clean one.
Two causes, both fixed:
- `git fetch --depth=1 origin $GITHUB_BASE_REF` fetches the base tip with no
shared history, so `base...HEAD` dies with "fatal: no merge base". Now
fetches --depth=50 with an explicit
+refs/heads/<base>:refs/remotes/origin/<base> refspec.
- The script only attempted three-dot. It now falls back to two-dot, which
needs no common ancestor, and reports which mode actually ran.
Verified against the exact failing condition — an orphan branch with no merge
base against origin/main, where `git diff origin/main...HEAD` returns
"fatal: no merge base": the guard now reports "diff check ran (two-dot ...)"
rather than skipping. NC1 re-confirmed: a leaf source change without a version
bump still fails, now accompanied by "diff check ran (three-dot ...)" proving
the check executed rather than being skipped into a pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f02759cd6d |
ci(guard): fail when a non-bundled @claude-flow/* leaf changes without a publish
Adds scripts/audit-leaf-package-publish.mjs, wired into v3-ci.yml's existing audit job (semver is already installed there). Only four @claude-flow/* components are bundled into the release tarballs — INTERNAL_RUNTIME_PACKAGES in stage-internal-runtime-bundles.mjs: security, codex, mcp, plugin-agent-federation. cli-core, neural, shared and memory resolve from the REGISTRY at whatever v3/@claude-flow/cli/package.json pins, so a merged source change in one of those does NOT ship with the three-package train: the release silently carries the last published copy. This has bitten twice. #3334 fixed RUFLO_INTELLIGENCE_MODE in @claude-flow/memory; v3.42.1 shipped the train only and a fresh consumer install had zero references to the fix (#3335). #3390's reasoningBank activation spanned cli/src/memory/memory-bridge.ts AND memory/src/controller-registry.ts — publishing only the train would have shipped the bridge half and left the embedder half unreachable, leaving reasoningBank exactly as inert as before. It was caught by hand and memory@3.0.0-alpha.25 went out alongside v3.42.5. What makes it easy to miss is that CI goes GREEN: the leaf's own package tests run against workspace source, so the change looks verified while being unshippable. The audit asserts: 1. every non-bundled leaf's workspace version is covered by the CLI's declared range (a miss means the release resolves to OLDER code than this repo's source); 2. diff mode — if a leaf's publishable source changed vs the base ref but its version did not, fail with remediation. It imports INTERNAL_RUNTIME_PACKAGES rather than restating it, so the bundled and non-bundled sets cannot drift apart; exits non-zero if it inspects zero leaves, so a broken classifier can't report clean; and reports a skipped diff check as NOT EXERCISED rather than as a pass. @claude-flow/shared is entered as a time-boxed accepted finding (expires 2026-10-21): the CLI pins 3.0.0-alpha.7 while the workspace is at 3.0.0-alpha.8, which IS published — stale, not unshippable. Raising the pin edits v3/@claude-flow/cli deps and so forces a pnpm-lock.yaml regen in the same change (#2540 -> #2552), which wants its own PR. An EXPIRED entry fails the guard rather than muting it. Negative-controlled, all four paths: - memory source changed, no bump -> FAIL (replays #3390) - same change with a version bump -> pass - test-only change in a leaf -> pass (not publishable) - waiver expiry backdated -> FAIL, "no longer honoured" Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b02d936935 |
fix(ci): print the failing assertion, not just the filename, on a ratchet failure (#3208)
The ratchet runs vitest with `--reporter=json --outputFile=…`, so nothing
vitest prints reaches the job log and the only record of an unexpected
failure is the file name. Recovering the assertion means downloading the
`test-results-ubuntu-latest` artifact — and a rerun replaces it, so after a
rerun attempt 1's assertion is gone for good.
The failure messages are already in the report this process parsed. Printing
them costs nothing, needs no new artifact, and survives a rerun, because job
logs are per-attempt.
`formatUnexpectedFailures()` is exported and pure so it unit-tests alongside
`evaluateTestReport()`. Every field it reads is optional by design: a
file-level abort ("No test suite found in file") carries `message` and no
assertions, and a report from another reporter version may carry neither. In
both cases the filename still prints, exactly as before, and the PASS path is
untouched.
Verified against a real failing run rather than a fixture: replaying run
35521170139's `vitest.json` through this script turns
+ v3/@claude-flow/cli/__tests__/memory-search-scores-3327.test.ts
into that same line plus
✗ memory_search score contract (#3327 Finding B) preserves standard fallback on import failure
AssertionError: expected [ { key: 'a', …(5) } ] to deeply equal [ { key: 'a', …(4) } ]
which is the line that would have identified the failure in #3208 without a
second run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KTedGmPKpEdArVBMMn3JaP
|
||
|
|
0a2df2bfd2 |
fix(metaharness): use the @metaharness/* and ruflo CLI already installed instead of npm-installing at tool-call time (#3366)
On a stock ruflo 3.42.4 install every metaharness_* MCP tool reached for the
npm registry at tool-call time, although the packages it needs are already on
disk. Two causes, in the same six helper files, which is why they ship
together: the helper PINS rejected what ruflo installs, and three code paths
never looked for an installed copy at all. Fixing only the pins leaves darwin
and the memory tools on npx; fixing only the lookups leaves them resolving a
pin that nothing on disk satisfies.
1. Pins (metaharness, @metaharness/darwin)
The metaharness_* tools resolve MetaHarness through the plugin helpers' own
tilde pins, not through package.json. Those pins were never bumped:
_harness.mjs stayed at ~0.3.0 after the CLI declared metaharness ^0.4.1
(
|
||
|
|
6f0ed71128 |
fix(release): spawn npm.cmd with shell:true in prepare-root-publish too (#3348)
Same EINVAL bug as #3346 (stage-internal-runtime-bundles.mjs's runBuild), one file over: prepare-root-publish.mjs's own npm.cmd spawnSync for building v3/@claude-flow/swarm and cli lacked shell:true, blocking the root claude-flow package's Windows publish the same way. Found immediately after #3346 while publishing 3.42.3 end to end. Single-string shell:true form, matching #3346's fix, to avoid Node's DEP0190 warning. Still constant literals, no injection surface. |
||
|
|
fcee45bc1d |
fix(release): spawn npm.cmd with shell:true on Windows (#3346)
* fix(release): spawn npm.cmd with shell:true on Windows
spawnSync('npm.cmd', ['run','build'], {stdio:'inherit'}) threw EINVAL on
Windows, blocking every @claude-flow/cli publish on this platform since
CreateProcess can't launch a .cmd directly and Node has refused to shell
out implicitly since CVE-2024-27980.
Found live while publishing 3.42.3: prepublishOnly -> prepare-publish.mjs
-> stageInternalRuntimeBundles() rebuilds the four internal runtime
packages (security/codex/mcp/plugin-agent-federation) and hit this on
the very first Windows publish attempt after upgrading past the Node
version where the implicit shell fallback still worked.
command/args here are constant literals ('npm.cmd', ['run','build']),
never hook-derived or user-controlled, so shell:true carries no
injection risk.
Co-Authored-By: RuFlo <ruv@ruv.net>
* fix(release): avoid DEP0190 warning in the npm.cmd shell:true fix
Single-string command form instead of args array — same shell:true
safety (still constant literals, nothing interpolated), no warning.
Co-Authored-By: RuFlo <ruv@ruv.net>
|
||
|
|
90889f4772 |
feat(sona): default the learning mode from RUFLO_INTELLIGENCE_MODE (#3334)
* feat(sona): default the learning mode from RUFLO_INTELLIGENCE_MODE The SONA learning mode was only settable per call (`config.mode`), defaulting to `balanced` otherwise. A deployment that wants every session to learn in a specific profile — a managed desktop fleet running `research` for +55% quality, or `edge` on constrained hosts — had no way to set that default short of threading `mode` through every call site or patching a generated helper. This reads `RUFLO_INTELLIGENCE_MODE` as the fleet-wide default in both the SONA config merge (`@claude-flow/integration` sona-adapter) and the memory learning bridge (`@claude-flow/memory`). Precedence is unchanged where a caller is explicit: `config.mode` (or `sonaMode`) still wins; the env only fills the former default slot; and an unset or unrecognised value returns undefined so the code falls through to `balanced` — a typo can never silently select a profile nobody asked for. Validated against the mode enum in each package. Co-Authored-By: claude-flow <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_01MUrJdDHhQ91hwpUBQhV4fC * fix(sona): resolve sonaMode per-instance in LearningBridge, not at module load learning-bridge.ts captured sonaModeFromEnv() into a module-scope DEFAULT_CONFIG constant, evaluated once on first import — so any RUFLO_INTELLIGENCE_MODE set later in the same process (including test setup) was permanently missed. sona-adapter.ts's equivalent (mergeConfig()) already re-reads the env var on every call; this brings LearningBridge's constructor in line with that behavior, preserving the same precedence (explicit config.sonaMode > env > default). Adds regression tests for both packages (sona-adapter.ts had none previously) and registers RUFLO_INTELLIGENCE_MODE in KNOWN_ESCAPE_HATCHES per the audit script's existing requirement for operator-knob env vars outside its SCAN_ROOTS. Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_01SjxUvZnyBdwtx4Y1jKAqt6 |
||
|
|
e036e35dd7 |
dream(memory): #3231 wire embedding near-dup detection into MemoryConsolidator.dedup() (evaluated, ACCEPT) (#3232)
* dream(memory): wire embedding near-duplicate detection into MemoryConsolidator.dedup()
MemoryConsolidator.dedup() only ever deduplicated entries via byte-exact
SHA-256 content hashing, even though every entry already carries a
computed .embedding and the adapter's HNSWIndex (already held in scope
for removal bookkeeping) is incrementally kept in sync. Paraphrases and
reformattings of the same memory were never caught.
Adds a second pass: for hash-pass survivors with an embedding, query the
already-populated HNSWIndex for cosine-similarity neighbors above a
configurable similarityThreshold (default 0.95, matching the unwired
domain-layer consolidator's own constant), and merge using the same
keeper-selection strategies (now factored into shared selectKeeper/
mergeGroup helpers). Guarded to cosine-metric indexes only; disabled via
similarityThreshold >= 1.
Evaluation: 471/472 passing in @claude-flow/memory (1 pre-existing,
unrelated, environmental failure); baseline (git-stash-isolated source)
fails exactly the 2 new discriminating tests and passes the rest,
confirming the fix is additive and non-regressive. tsc --noEmit clean.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG
* dream(memory): fix-forward — converge near-dup pass to a fixed point in one dedup() call
Independent adversarial critique (STEP 10) reproduced a real, bounded gap
in the prior commit: a near-duplicate cluster larger than
NEAR_DUP_SEARCH_K (8) split into multiple leftover sub-group survivors
within a single dedup() call, because each round's `consumed` bookkeeping
permanently excluded a group's keeper from further matching even though
it was still fully present in the index. 15 pairwise-identical-embedding
entries collapsed to 2 survivors (merged: 13) instead of 1 (merged: 14).
It self-healed across repeated runAll() calls (background timer /
nightlyLearner), so this was never a permanent-data-loss bug, but a
single manual dedup() call under-converged relative to its documented
"collapse duplicates" contract.
Fix: loop pass 2 to a fixed point (re-scan until a full round produces
zero merges) instead of a single scan. Always terminates — entries
strictly decrease each round that merges anything. `groups` now counts
merge operations across all rounds, which can exceed the number of
underlying duplicate clusters when one needed more than one round;
documented in the method's doc comment.
Added a discriminating regression test reproducing the critic's exact
15-entry scenario; confirmed it fails against the pre-fix single-round
code (merged: 13) and passes against this fix (merged: 14, single
survivor). Full package suite: 472/473 passing (1 pre-existing,
unrelated, environmental failure, unchanged from before this commit).
tsc --noEmit clean.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG
* dream(memory): 2026-09-08 research gist — embedding near-dup consolidation
Full SOTA report: 5-role parallel research fan-out, ledger check with
GitHub-verified fates for the last 14 rows (3 merged since 09-03, one PR
now crossing the 14-day stale threshold for the first time this cycle),
competitor comparison (CrewAI/Mem0/Zep/LangMem/Qdrant/SemDedup), plugins
and automation scan findings, adversarial critique with a real bug found
and fixed, and witness stamp.
No gist-creation MCP tool available in this session (GitHub MCP tools
cover issues/PRs/repos, not gists) — committing to docs/dream-cycle/
instead, matching every dream-cycle night since 2026-08-14.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG
* dream(memory): #3231 update LEDGER.md for 2026-09-08
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG
* dream(memory): #3231 empirical receipts for dedup() near-dup pass (response to review)
Response to PR #3232's INCONCLUSIVE review, which asked for a threshold
contract rather than a smoke fixture: a contamination-checked embedding
corpus with pinned model, precision/recall including hard negatives and
cross-namespace cases, baseline vs. candidate latency/memory at realistic
cardinality, and deterministic replay + unchanged exact-dedup behavior.
Real pinned-model corpus: attempted directly against both ONNX backends
this monorepo ships. Both fail for reasons unrelated to this candidate,
reproduced on an unmodified checkout before writing anything:
- @huggingface/transformers@4.2.0: transformers.node.mjs does
`import { Tensor } from "onnxruntime-common"` against a CJS module — a
real ESM/CJS interop break under this sandbox's Node version.
- @xenova/transformers@2.17.0: pulls in sharp@0.32.6, whose native
binding (sharp-linux-x64.node) isn't present for this platform — the
same failure independently reproduced moments later by this package's
own nightlyLearner test, which falls back to mock embeddings.
Neither is fixable within this PR's scope without touching unrelated,
pre-existing native/module-resolution infrastructure. Documented in
detail at the top of the new test file rather than silently worked
around or omitted.
What's delivered instead, in
v3/@claude-flow/memory/src/consolidator-embedding-benchmark.test.ts:
- A deterministic synthetic corpus where pairwise cosine similarity is
constructed EXACTLY via vectorAtSimilarity() (Gram-Schmidt against a
random orthogonal vector), not measured after the fact — a rigorous
instrument for threshold mechanics specifically, disclosed as not
validating real-world semantic accuracy.
- Precision/recall/false-merge-rate across a 5-point threshold sweep,
every point at a deliberate margin from every corpus similarity value
after an exact-boundary collision (sweep value == corpus value) proved
float32-jitter-sensitive while writing this — documented as a finding,
not silently patched around.
- Namespace isolation: proved cross-namespace near-dup merging matches
cross-namespace hash-exact merging exactly (both pre-existing,
unscoped-by-namespace behavior — this candidate doesn't change it).
- Determinism: 3 independent runs, compared by deterministic entry KEY
(not raw id, which is randomly generated per store() and was a second
bug this file's own first draft had to fix).
- Latency/memory at N=5000 (matching this repo's own HNSW-benchmark
convention): initially measured 9.8s for the candidate pass, which
traced to an unset `ef` search parameter defaulting to efConstruction
(200) — every dedup() search traversed a 200-candidate list for a
duplicate-detection task that only needs very-close neighbors. Fixed
by passing an explicit NEAR_DUP_SEARCH_EF=32 in consolidator.ts,
re-measured at 3.7s (~2.7x), precision/recall unchanged (still exact).
Real numbers logged in the test output either way, not asserted away.
Full package suite: 478/479 passing (472 pre-existing + 7 new; same 1
pre-existing unrelated environmental failure as every prior commit on
this branch). tsc --noEmit clean.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG
* dream(memory): #3231 add new benchmark file to CI test-ratchet baseline
CI on a2e99cf failed with "1 unexpected failing file(s): +
v3/@claude-flow/memory/src/consolidator-embedding-benchmark.test.ts".
Root-caused via the vitest JSON report artifact (test-results-ubuntu-latest,
downloaded and parsed directly, not guessed): a suite-level collection
failure, "Cannot find package '@claude-flow/security' imported from
'.../v3/@claude-flow/memory/src/agentdb-retrieval-guard.ts'" — the same
unbuilt-sibling-package gap already accepted in scripts/ci-test-baseline.txt
for 9 other @claude-flow/memory test files, including consolidator.test.ts
itself (line 95). Every file in this package fails identically in this CI
checkout state; my new file just wasn't in the baseline yet because it's
new. Not a regression this PR introduced — confirmed by the failure being
a pre-import-time module-resolution error, not a test assertion in the
new file's own logic (which passes locally in a checkout with @claude-flow/
memory's dependencies actually built: see the 6/6 local run in the prior
commit's message).
Fix: add the new file to the baseline, alongside its 9 memory-package
siblings already there, matching this repo's own established convention
for this exact gap rather than working around it some other way.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01H4kfasS4XkYaSGYXFLSSEG
* dream(memory): #3231 baseline 2 more federation/gateway plugin test files
CI on
|
||
|
|
1bd734e272 |
dream(security): #3102 bind ADR-377 caller-identity verification into live authorizeMcpTool chokepoint (evaluated, ACCEPT-with-caveats) (#3103)
* dream(security): #TBD bind ADR-377 caller-identity verification into live authorizeMcpTool chokepoint (evaluated, ACCEPT-with-caveats)
authorizeMcpTool() trusted the plain, unsigned CLAUDE_FLOW_PRINCIPAL_ID env
var as caller identity with zero verification (OWASP ASI07-class gap).
ADR-377 Phase 3 already implemented Ed25519 issueInvocationToken/
verifyInvocationToken but never wired it into a live dispatch path.
Binds the two: DualModeOrchestrator mints a per-worker signed
InvocationToken at spawn (worker-lifetime TTL, wildcard tool scope --
disclosed reduction from the primitive's original per-call design);
authorizeMcpTool verifies it before trusting identity, failing closed on
missing/forged/expired tokens. Off by default (CLAUDE_FLOW_MCP_CALLER_AUTH).
6 new deterministic Vitest scenarios plus an independent adversarial
critique (fresh session, verdict CONFIRMED-SAFE-WITH-CAVEATS) are recorded
in docs/dream-cycle/dream-gist-2026-08-26.md. Darwin skipped (scope
mismatch -- binary crypto gate, no tunable parameter space).
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_015bfVASxkNXKpZbMSLhVU1L
* fix(ci): register ADR-377 Phase 3 caller-identity env vars as known escape hatches
CLAUDE_FLOW_MCP_CALLER_PUBKEY and CLAUDE_FLOW_MCP_INVOCATION_TOKEN
(DualModeOrchestrator's per-worker Ed25519 credential pair, read by
resolveMcpCallerIdentity()) weren't registered in
audit-env-var-precedence.mjs's KNOWN_ESCAPE_HATCHES, so the audit flagged
CLAUDE_FLOW_MCP_CALLER_PUBKEY as a real violation (exit 1). Same no-CLI-flag
reasoning as the existing CLAUDE_FLOW_PRINCIPAL_ID entry: these are
credentials minted into a spawned worker's own environment, not a value
any caller should be able to select via a flag.
Verified: audit-env-var-precedence.mjs now exits 0. policy-runtime.test.ts
still 18/18.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01RYP825phHMjPRmmjkcAWpq
* fix(ci): regenerate root package-lock.json for codex's new @claude-flow/security dep
The CI Test Suite job runs plain `npm ci` at repo root (no v3 pnpm
install step), so it resolves workspace deps strictly from
package-lock.json — not from v3/pnpm-lock.yaml, which this PR had
already updated. codex/package.json's new @claude-flow/security
dependency was missing from package-lock.json's codex entry, so CI's
npm ci never linked it and every dual-mode test importing
orchestrator.ts failed with "Cannot find package '@claude-flow/security'".
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01RYP825phHMjPRmmjkcAWpq
* fix(ci): build @claude-flow/security before Test Suite runs it
Complements
|
||
|
|
8008bc3562 |
fix(federation): verified event identity must win over publisher-controlled content (#3322)
* fix(federation): verified event identity must win over publisher-controlled content fetchRecent/fetchManyOn/fetchChannel spread the parsed message content after the signature-verified id/pubkey/created_at, so a publisher could put those keys in their own content and overwrite their verified identity downstream (reduceClaims trusts .pubkey for claim release/handoff authorization). Reorder so verified fields always win. Also in x-federation-channels.ts: reqEvents() never called verifyEvent() at all, so channel-read/grant-accept events were trusted unsigned; add the check. Same content-override reorder in x_federation_channel_read. fix(hooks): escape argv for the Windows shell:true npm-shim path ruflo-hook.cjs's invokeHook() passes hook-derived values (command text, file paths) through cmd.exe via shell:true to resolve npm's .cmd/.ps1 shims. Node does no escaping in that mode, so a value containing a cmd.exe metacharacter could be reinterpreted as a separate command or redirection instead of reaching the CLI as data. Add escapeCmdArg() (qntm.org/cmd algorithm, same one cross-spawn uses) and apply it to every argv element on that path. No native Windows/cmd.exe available to execute this end-to-end in this environment — verified at the string-transform level only (8 unit tests). Real Windows execution remains a required follow-up gate. Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_01RYP825phHMjPRmmjkcAWpq * fix(ci): register new hook test env var and test file with CI guards Follow-up to the previous commit's escape-cmd-arg.test.cjs: - Register RUFLO_HOOK_UNIT_TEST in audit-env-var-precedence.mjs's KNOWN_ESCAPE_HATCHES, matching the existing RUFLO_HOOK_CLI_OVERRIDE/ RUFLO_HOOK_DEBUG_STDOUT entries for the same file. - Register plugins/ruflo-core/scripts/escape-cmd-arg.test.cjs in ci-test-baseline.txt, matching the sibling mcp-launch.test.cjs entry — both are node:test-based .cjs files that vitest's own JSON reporter marks "failed" (it doesn't recognize the node:test API) even though every test inside genuinely passes under `node --test`. Verified locally: node scripts/audit-env-var-precedence.mjs now passes (exit 0). A scoped `vitest run` + `ci-test-ratchet.mjs --report` against just these two files confirms escape-cmd-arg.test.cjs is no longer an unexpected failure (the one remaining flagged file is a stale local worktree artifact under .claude/worktrees/, not part of a clean CI checkout). Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_01RYP825phHMjPRmmjkcAWpq |
||
|
|
8485302f10 |
feat(x-gateway): 0.6.1 — bound Seraphina by budget instead of an admin token (#3275)
* feat(x-gateway): 0.6.1 — Seraphina bounded by budget, not by an admin token Squashes what is actually running on x.ruv.io so git and production agree. Seraphina was gated alongside the write tools, but it is not like them: it reads the roster, the claims board and recent messages, asks a model, and returns advice. It writes nothing and carries no authority. The gate answered the wrong question — the exposure is model spend, and spend is bounded with a budget, not a password. The practical cost was worse than a wrong abstraction. A browser cannot hold a bearer secret, so a published UI was either locked out or tempted to ship RUFLO_ADMIN_TOKEN to the client — the same token that mints invites and publishes as the gateway identity. The safe-looking option was the catastrophic one. Seraphina now answers with no token, bounded by a shared daily cap and a per-client hourly cap, and anonymous callers cannot select the high or ultra tiers. An admin token lifts both. Every write tool keeps its gate; verified live that claims_issue, federation_invite_mint, federation_publish, federation_admit and channel_publish all still refuse an anonymous caller. Not included, deliberately: a tool to relay member-signed events through the gateway. It was built, tested against the live relay, and removed. buzz-relay refuses any EVENT whose pubkey differs from the NIP-42 authenticated connection, so a gateway cannot publish on another identity's behalf — and it does not need to. Members already publish as themselves through the wss://x.ruv.io proxy, verified end to end. Shipping a tool that always fails would be worse than none. Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67 * feat(x-gateway): serve onboarding guidance, and deliberately not key generation The registry said "generate a Nostr keypair" and stopped, and it never named the two identities in play. That gap was not theoretical: a capable agent inspected the surface, found only gateway-signed publishing, and concluded ordinary members could not publish from a dashboard at all. They can. It just was not written anywhere they could reach. federation_onboarding and ruv://federation/onboarding now answer it, open, no token. The guide leads with the distinction that causes the confusion — your key signs for you, the gateway's key signs for the service, and the gated tools are gated precisely because they speak as the service. It carries the five steps, the four things never to do, and the three traps this service has actually taught us: sign the NIP-42 relay tag with the canonical URL even when proxied, the relay binds publishing to the authenticated connection, and the channel tag is `c` not `h`. It does NOT generate keys, and that is the point rather than an omission. A service that mints your keypair and returns the secret has seen your secret, and becomes custodian of every identity it "helped" — the same custody mistake as putting an admin token in a browser, inverted. The guide ships the code so the caller runs it locally and the key never crosses the wire. One test note worth keeping: the first version of the safety assertion regex- matched for "send your secret" and failed on the guide's own "Never send your secret key" line. It is now structural — no field may be named for secret material — because a string search cannot tell an instruction from its negation. Verified live on 0.7.0 (ruflo-x-gateway-00015-qj7) with an anonymous call. 20/20. Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67 * fix(ci): register the gateway's credential and budget knobs, and skip test/ like tests/ The ADR-125 precedence audit failed the gateway PR, and it was right to look — but for two reasons that are both about the audit, not the code. RUFLO_ADMIN_TOKEN was not registered as an escape hatch even though its sibling RUFLO_X_ADMIN_TOKEN is, with the same reasoning: a secret must never be a CLI flag, where it lands in shell history and process listings. The gateway is a long-running service with no typed command surface at all, so there is no invocation to attach a flag to. Same for the two Seraphina budget knobs added in #3275. The other half was a naming gap. SKIP_DIRS already excludes `tests` and `__tests__` because the audit is about production precedence, not test setup — but the flagged lines were in `plugins/ruflo-x-gateway/test/`, singular, which was not in the set. Two `process.env.X = 'test-admin-token'` assignments inside a test fixture were reported as undeclared production reads. Verified: the audit now exits clean on this branch. Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67 |
||
|
|
dbf450a927 |
feat(federation): ruflo CLI + MCP integration for x.ruv.io, and Seraphina (swarm queen) (#3256)
* feat(federation): ruflo CLI + MCP integration for x.ruv.io, and Seraphina (swarm queen)
Integrate the open swarm federation into ruflo itself, and add Seraphina —
a primary-coordinator / swarm-queen guidance tool in the style of the ruOS
assistant terminal (invoked from a terminal or any MCP client).
MCP tools (src/mcp-tools/x-federation-tools.ts), registered in the barrel and
the in-process registry so `callMCPTool` and the MCP server both see them:
x_federation_sync / roster / claims / registry (open reads)
x_federation_publish / invite_mint / admit (gateway-identity writes,
require RUFLO_X_ADMIN_TOKEN; fail closed with no network call)
Each description follows ADR-112 ("Use when … wrong because …").
Seraphina (src/mcp-tools/seraphina-tools.ts): `seraphina_guidance { goal }`
gathers the live roster, claims board and recent messages from the x.ruv.io
gateway, compacts them (dedupe by from|type, cap 15 — a cheap tier drowns in
repeated PeerHellos), and asks the cognitum meta-llm gateway
(https://api.cognitum.one/v1/messages, model cognitum-auto by default with a
tier override) using a queen system prompt that treats message content as
data, respects one-owner-per-resource claims, and returns
{ guidance, proposals[], risks[] }. Proposals are advisory. JSON is extracted
by slicing the outermost object so fenced/prose-wrapped answers still parse.
Key from SERAPHINA_METALLM_KEY (GCP secret seraphina-metallm-api-key).
CLI: `ruflo federation sync|roster|claims|registry|invite|admit|publish`
(src/commands/federation.ts), thin over the same tools.
Skill: .agents/skills/open-federation/SKILL.md documents CLI, tools,
Seraphina, the claims rules, the canonical NIP-42 relay-tag gotcha, and
onboarding.
Tests: 11 (x-federation: RPC shaping, SSE parse, resource mapping, fail-closed
admin gating, isError surfacing; seraphina: registration, fail-closed key,
context gathering + queen prompt + x-api-key, tier override, fenced-JSON
extraction). Project typechecks at 0 errors.
Verified live: Seraphina against the real 3-node federation returned 3
structured proposals + 4 risks via cognitum-auto (routed to a cheap tier).
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
* fix(federation): ADR-125 env-var precedence — URL config takes a flag/arg, credentials registered env-only
The env-var-precedence audit correctly flagged the new process.env reads.
Config URLs now have a precedence path: `ruflo federation --gateway` feeds a
`gatewayUrl` tool arg (and Seraphina takes `metaLlmUrl`/`gatewayUrl`) that
wins over RUFLO_X_GATEWAY_URL / SERAPHINA_METALLM_URL, documented per ADR-125.
Credentials (RUFLO_X_ADMIN_TOKEN, SERAPHINA_METALLM_KEY) are registered as
env-only escape hatches with rationale: a secret must never be a CLI flag.
+1 test (arg precedence, arg not forwarded upstream). Audit passes locally.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
* feat(federation): `ruflo federation join --code` — self-service invite access with your own key
The user-facing path into the open swarm. Decentralized by design: the user
generates/holds THEIR OWN Nostr key (~/.ruflo/nostr.key, 0600), redeems an
invite code with a NIP-98-signed claim directly against the relay (no admin in
the loop), proves membership via NIP-42, and can then publish as themselves.
The gateway never signs for a user.
- src/mcp-tools/x-federation-join.ts: x_federation_join { code, relayHttp?,
relayWs?, keyFile? } — validates the code shape before any network call,
loads/creates the key, NIP-98 claim, NIP-42 verify, returns pubkey + role.
ADR-125: args take precedence over RUFLO_X_RELAY_HTTP / RUFLO_X_RELAY_WS /
RUFLO_NOSTR_KEY_FILE.
- CLI: `ruflo federation join --code v2.…`
- nostr-tools added as an optionalDependency (secp256k1/Schnorr is not in
node:crypto); the tool degrades with an install hint when absent. pnpm
lockfile regenerated in this PR (frozen-lockfile CI).
- 4 tests: 0600 key create/reuse, valid NIP-98 header (kind 27235, verifies,
bound to url+method+payload hash), malformed code rejected pre-network,
ADR-112 description. Suite: 16/16; tsc 0; env-var audit passes.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
* fix(federation): default relay -> wss://relay.ruv.io; join test tolerates missing nostr-tools
- x-federation-join defaults now point at the canonical relay.ruv.io host (Cloud Run
domain mapping for buzz-relay); the raw run.app host stays routable for old clients.
- Validate the invite-code shape before the optional-dependency check so a bad code
fails fast whether or not nostr-tools is installed.
- The root Test Suite runs the CLI tests via root npm ci, which never installs the
CLI's optionalDependencies: the crypto cases are it.skipIf(!nostr-tools) so the
file no longer trips the CI test ratchet.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
|
||
|
|
6396ff636c |
feat(x-gateway): open Nostr swarm federation gateway — MCP + ruv:// + ws proxy (deployed at x.ruv.io) (#3255)
* feat(x-gateway): MCP + ruv:// gateway for open Nostr swarm federation
New plugin `ruflo-x-gateway` — the service behind x.ruv.io. An MCP server
(Streamable HTTP at /mcp) that exposes ruflo swarm FEDERATION and CLAIMS over
an open, membership-gated, SIGNED Nostr relay (buzz-relay).
Why Nostr: every coordination message is a signed Nostr event, so authorship is
cryptographically verifiable and anyone the relay admits can participate — open
but secure. The relay gates membership + NIP-42 auth; no pre-pinning needed.
Surface:
- Tools: federation_identity / federation_join / federation_publish /
federation_sync, claims_issue / claims_release / claims_status.
- Resources: ruv://federation/registry, ruv://swarm/roster, ruv://claims/board.
- src/nostr-federation.mjs: NIP-42 authenticated connect, signed publish, and
verified fetch of #t=ruflo-swarm coordination events.
- src/server.mjs: node http server routing /, /health, /mcp (stateless
Streamable HTTP), plus the ruv:// resources.
- Dockerfile for Cloud Run; persistent Nostr identity at /data (0600).
Smoke-tested locally: server starts, /health + / respond, POST /mcp tools/list
returns the tool set over SSE. Federation tools require relay membership (by
design) — deployment wires membership + DNS (x.ruv.io) as follow-up.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
* feat(x-gateway): stable identity via RUFLO_NOSTR_KEY_HEX (GCP secret) + read-only-FS fallback
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
* feat(x-gateway): v0.2.0 — ws proxy, admin-gated writes, invite/admit tools, tests
Security model made explicit: /mcp is public, so every tool that writes
using the GATEWAY's own identity (join, publish, claims_issue/release,
invite_mint, admit) now requires `adminToken`, checked constant-time and
fail-closed (no configured token => all writes denied). Reads and ruv://
resources stay open. Users publish with THEIR OWN keys via invite->claim.
- src/ws-proxy.mjs: transparent WebSocket proxy so wss://x.ruv.io (and
/relay) fronts the Nostr relay; 500-conn cap, 502/503 on failure.
- src/relay-admin.mjs: mintInvite (NIP-98 POST /api/invites) and
admitMember (NIP-43 kind 9030) — the gateway holds relay admin role, so
the owner key never leaves GCP.
- src/security.mjs: per-IP token-bucket rate limit (60/min), 256KB body
cap enforced before buffering, security headers, timingSafeEqual admin.
- src/claims.mjs: owner-per-resource reducer extracted for testing.
- src/server.mjs: createGateway() factory (testable), stateless MCP.
- test/gateway.test.mjs: 8 tests — claims rules, gating, rate limit, body
cap, NIP-42 against a mock relay (verifies the signed challenge), routes.
Verified live before this commit: gateway pubkey admitted + promoted to
relay admin; invite minted and a fresh key self-joined via claim and
passed NIP-42 auth; deployed gateway publishes/reads over the relay.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
* fix(x-gateway): surface canonicalRelay — relay verifies NIP-42 relay tag strictly
Empirical: auth through the wss://x.ruv.io proxy is ACCEPTED when the client
signs relay=<canonical relay URL> and REJECTED (verification failed) when it
signs relay=wss://x.ruv.io. Expose canonicalRelay + authNote at GET / and in
ruv://federation/registry so clients sign the right tag.
Co-Authored-By: RuFlo <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_013u4pmL9ZUAXb6usVQgNo67
* ci: retrigger test-suite (ratchet flake; main green at
|
||
|
|
13cbd697db | fix(release): enforce CI and idle assignment safety | ||
|
|
21f7c0adb8 | fix(release): align bundled runtime metadata | ||
|
|
2ec82b0cd1 | chore(release): prepare stable 3.38.10 train | ||
|
|
1e24bb8878 |
fix(agent): propagate explicit provider/model config into agent execution (#2962) (#3007)
* fix(agent): propagate explicit provider/model config into agent execution (#2962) `providers configure` and `agent spawn --provider/--model` persisted user intent but the execution path (callAnthropicMessages, executeAgentTask, determineAgentModel) only ever consulted env vars and a 5-alias model list, silently discarding both. A local Ollama/OpenRouter setup with explicit config would either fail closed or fall back to whatever provider the env vars happened to select. - determineAgentModel(): treat any non-alias config.model string as an explicit selection via the existing modelId fast-path, instead of falling through to task-based routing / agent-type defaults. - agent_spawn: forward config.provider to the stored (and returned) agent record when it's an unambiguous explicit choice ('ollama'/'openrouter'; 'anthropic' is excluded since the CLI silently defaults to it when --provider isn't passed). - callAnthropicMessages(): accept an optional provider param and consult the persisted `agents.providers` config for baseUrl/apiKey/model when env vars are absent. A self-hosted Ollama baseUrl no longer requires the undocumented OLLAMA_API_KEY='local' sentinel. - executeAgentTask(): forward agent.provider into the first dispatch call. Precedence: explicit per-agent flag > env vars > persisted config > key-presence inference (unchanged, last resort). Co-Authored-By: RuFlo <ruv@ruv.net> * fix(ci): update #2042 smoke to match the widened OpenRouter branch shape The #2042 regression smoke statically matched the literal token sequence `useOpenRouter && openrouterKey`. #2962 widened that condition to `if (useOpenRouter) { const apiKey = openrouterKey || persistedOpenRouter?.apiKey; ... }` so a persisted `agents.providers` config can supply the key when no env var is set — the smoke's ordering guarantee (OpenRouter branch reachable before the Anthropic-key early-return) still holds, only the literal pattern needed updating. Co-Authored-By: RuFlo <ruv@ruv.net> |
||
|
|
83e536396f |
fix(scaffold): add dead-reference guard + deterministic CLI remap (ADR-382 Part C) (#2974)
Adds scripts/smoke-init-scaffold-references.mjs (ADR-382 Part C, #2971): four static assertions over v3/@claude-flow/cli/.claude/** and the plugin/marketplace surface, deriving the live MCP tool set and canonical CLI form from source rather than hand-maintaining them. 1. dead `npx claude-flow` (bare) invocation 2. dead `mcp__claude-flow__<tool>` references not in the live registry 3. plugin .mcp.json launches with no local-bin-first resolver (Part A regression guard) 4. plugins/* directories missing from .claude-plugin/marketplace.json Ships warn-only (no --strict) with a documented --strict flag to flip once backlogs clear. Wired into .github/workflows/v3-ci.yml as a new job, gated on plugins/*/.mcp.json, .claude-plugin/marketplace.json, and the script itself (v3/@claude-flow/cli/.claude/** was already a trigger path). Deterministic remap applied: 701 occurrences of bare `npx claude-flow` across 142 files -> `npx @claude-flow/cli@latest` (mechanical, 1:1, verified by rerunning the guard: check 1 701 -> 0, check 2 unchanged at 410). The 386 legacy `npx claude-flow@alpha` / `@v3alpha` occurrences elsewhere in the same tree were left untouched (still-maintained dist-tags). NOT remapped in this PR (tracked follow-up, per ADR-382 Part C's own guidance not to guess): 410 dead `mcp__claude-flow__<tool>` occurrences across 98 files, 54 distinct dead tool names (memory_usage: 53 files, task_orchestrate: 41, sparc_mode: 34, swarm_monitor: 16, plus 50 more at lower counts). Each requires per-call-site judgment about the surrounding example's intent (store vs retrieve vs list, 1-line vs restructured 2-line orchestration calls) that a bulk regex pass cannot make safely. Checks 3-4 are expected-red until ADR-382 Part A merges (plugin/ruflo-core .mcp.json resolver + the 3 missing marketplace entries: ruflo-agntcy, ruflo-bbs-federation, ruflo-business-pods). |
||
|
|
f35c545fbe |
feat(metaharness): pull in @metaharness/turn-credit + fix stale router/darwin pins (#2958)
* feat(metaharness): pull in @metaharness/turn-credit + fix stale router/darwin pins (post metaharness#176) metaharness#176 shipped @metaharness/turn-credit (ADR-248, recursive turn-level credit assignment) and bumped darwin to 0.9.0 (ADR-249 signal seams) + router to 0.4.0 (calibration module). This brings ruflo's dependency contract back in sync and makes the new package available: - Add @metaharness/turn-credit ~0.1.0 as a new optionalDependency, following the same "must be installable, not peer-only" pattern darwin/flywheel/radio already use (ADR-150) — dependency-free, 64.9K unpacked, zero lifecycle scripts, same profile as the other three. - Bump @metaharness/darwin ~0.8.3 -> ~0.9.0 and the paired MH_DARWIN_PIN constant in distill-oracle.ts (was already out of range: tilde only absorbs patches, and 0.9.0 is a minor bump). - Bump @metaharness/router peer range ^0.3.2 -> ^0.4.0 (still deliberately peer-only + triple-gated behind CLAUDE_FLOW_ROUTER_NEURAL=1, per neural-router.ts — that design choice is unchanged, only the stale range is fixed) and update the matching manual-install hint in neural.ts. - scripts/check-metaharness-pins.mjs + scripts/metaharness-clean-install-test.mjs: add turn-credit to the watched/contract-checked package list. - .github/workflows/no-cli-optdep-bloat-2561.yml: CLI_MAX 10 -> 13. The prior bump (PR #2956) left zero slack (budget == count exactly), which a code review flagged as a latent trap — the very next unrelated optional dep would trip this guard. This bump leaves real headroom (11 declared today, budget 13) instead of repeating that mistake. - Also fixes a live ReferenceError in distill-oracle.test.ts (MH_DARWIN_PIN used but never imported) — a gap from the prior release's test fix that somehow didn't surface in that PR's CI; caught here while touching the same file. Co-Authored-By: RuFlo <ruv@ruv.net> * fix: regenerate v3/pnpm-lock.yaml — was out of sync with package.json edits The previous commit edited v3/@claude-flow/cli/package.json directly (darwin/router/turn-credit pin changes) without regenerating the pnpm workspace lockfile, so every CI job running `pnpm install --frozen-lockfile` failed immediately with ERR_PNPM_OUTDATED_LOCKFILE — cascading into every downstream smoke/test job that depends on that install step. Regenerated with pnpm@8.15.9 (matching CI's pinned version) in an isolated worktree to avoid the lockfile-version drift a newer local pnpm would introduce. Co-Authored-By: RuFlo <ruv@ruv.net> |
||
|
|
0b3cfb77d6 |
MetaHarness hardening: repair the dependency contract + strict sequential promotion evidence (#2956)
* feat(metaharness): repair dependency contract + strict sequential promotion evidence Item 1 — dependency & compatibility repair (the contract was silently broken): - @metaharness/darwin ^0.8.3, @metaharness/flywheel ^0.1.10, and @metaharness/radio ^0.1.0 are now explicit optionalDependencies of @claude-flow/cli (optional PEER deps are never auto-installed, so a clean ruflo install shipped with zero MetaHarness packages on disk) - check-metaharness-pins.mjs now searches dependencies, optionalDependencies, AND peerDependencies; UNDECLARED and PEER-ONLY (for installable pins) are fatal drift instead of silently passing; radio added to the watch list; --require-installed asserts the real published symbol contracts (RefineMutator, withSequentialEvidence, RadioBus, ...) - new scripts/metaharness-clean-install-test.mjs: installs the declared ranges into a pristine temp dir and asserts every advertised export; wired as a MANDATORY clean-install job in metaharness-ci.yml - doctor: new "MetaHarness declared packages" check FAILS (not warns) when a declared optional dep does not resolve at runtime; --component metaharness now runs upstream + declared-deps + integration checks - MH_DARWIN_PIN bumped 0.8.0 → 0.8.3 in lock-step with the declared range Item 2 — strict promotion evidence at the transaction authority: - receipts now carry task-level pairedOutcomes (taskId + per-task baseline/ candidate scores) behind heldOutDeltas; verifyFlywheelReceipt refuses rows that cannot reproduce their aggregate; evaluateFlywheelCandidate populates them; pre-existing receipts still verify byte-identically - new flywheel-sequential-evidence.ts: anytime-valid e-process over discordant paired outcomes (testing-by-betting) composed with per-candidate alpha allocation (alpha_k = alpha_total * 6/(pi^2 k^2)), so the family-wise false-promotion probability across an ADAPTIVE candidate stream is bounded by alpha_total = 5% - promoteFlywheelCandidate requires paired evidence by DEFAULT — aggregate- only receipts are refused, never silently downgraded (the upstream withSequentialEvidence fallback hole); explicit --allow-aggregate-evidence / allowAggregateEvidence escape hatch for pre-upgrade receipts; alpha spend is persisted per receiptId in the transaction state (looking spends alpha; retries reuse their index) - acceptance test: 1,000 null-improvement streams (20 adaptive candidates x 40 worst-case discordant pairs) — measured family-wise false promotion 0.6%, within the <=5% budget Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy * chore: gitignore node-compile-cache build artifact Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy * fix(ci): satisfy #2561 guard, dep-overlap audit, and smoke contract for metaharness optional deps Three CI failures from the dependency-contract repair, three fixes: - tilde-pin @metaharness/{darwin,flywheel,radio} (~0.8.3 / ~0.1.10 / ~0.1.0) per the ADR-150 anti-caret rule enforced by smoke step 17z73 — upstream is unstable, tilde absorbs patches but never minors - remove the three from peerDependencies/peerDependenciesMeta: an entry in BOTH optionalDependencies and peerDependencies crashes npm 11.x arborist on dedupe (#1147/#2018 dep-overlap audit); optionalDependencies alone is the declaration that actually installs - update the #2561 cold-startup guard per its own escape clause: budget 8 → 10 and drop @metaharness/darwin from the forbidden list — that entry dated from the pre-0.8 heavy-tree era; darwin@0.8.3 is 1.8M unpacked with ZERO dependencies and no lifecycle scripts (all three packages combined: 2.3M, 753ms cold install into an empty dir, measured 2026-08-10; the clean-install CI job re-verifies on every pin change) - smoke.sh 17h now accepts the componentMap ARRAY form ('metaharness': [checkMetaharness, ...]) introduced with the declared-packages doctor check Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy * fix(ci): remove apostrophe that broke the single-quoted guard script The #2561 guard runs as `node -e '...'` inside bash; an apostrophe in a comment ("guard's") terminated the quoted string and bash tried to execute the next // comment line (exit 126, '//: Is a directory'). Verified by executing the extracted run block end-to-end: all three checks OK, exit 0. Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy * feat(flywheel): ADR-381 — alpha-stream governance, reset epochs, accept/v2+seq Completes the sequential-evidence governance PR #2956 opened: ADR-381 (new) records the decisions: - the ADR-322 promotion ledger IS the alpha stream (per project root) — receipt lineageIds default to fresh UUIDs and cannot scope the control - evidence epochs: resetSequentialEvidence({confirm, reason}) archives the spend into an append-only sequentialResets audit trail, EXPIRES every outstanding evaluated receipt (fresh-data enforcement — old-epoch evidence cannot be replayed against a reopened budget), and increments evidenceEpoch; surfaced as `flywheel evidence-reset --reason … --confirm` (CLI) and the metaharness_flywheel MCP op, both behind the same policy gate as promotion - accept/v2+seq: the ADR-176 generations loop now decides each bundle with a third conjunct — the e-process over the bundle's own embedded per-task holdout at alpha_k for its position in the attempts stream — with the full verdict recorded in the bundle and independently replayed by verifyReceiptBundle (v1 bundles keep v1 semantics; versions pin per bundle); flywheelStatus surfaces next test index / threshold / minimum pairs / remaining budget so exhaustion reads as plateau, not mystery - two-layer pre-flight: evaluateFlywheelCandidate annotates (never blocks) when the promotion holdout cannot clear the next threshold on a perfect sweep; promoteFlywheelCandidate refuses size-inviable receipts BEFORE allocating an alpha index — sample size is ancillary, so the refusal looks at no evidence and spends no budget Tests: reset semantics (archive/expire/epoch/fresh-promote), alpha-free size refusal vs alpha-spending e-process refusal, v2 promote/reject/replay including tamper + index-shopping detection, v1 compat, helper math anchors (min pairs = 9 at test 1), budget monotonicity. 82 tests green across the eight affected suites; tsc build clean. Co-Authored-By: RuFlo <ruv@ruv.net> Claude-Session: https://claude.ai/code/session_01FKWLVWewYaWmcUg2q8Y9wy * fix(flywheel): close 3 concurrency/epoch gaps in ADR-381 sequential evidence Code review of this PR found three real correctness bugs that can silently violate the <=5% family-wise false-promotion guarantee that is the PR's own core deliverable: - flywheel-transaction.ts: resetSequentialEvidence's "fresh data only" guarantee only expired receipts already registered at reset time; a receipt whose evidence predates a reset but is registered afterward was silently admitted into the new, cheaper epoch (index shopping). Now tracks evidenceEpochStartedAt and promoteFlywheelCandidate refuses any receipt whose payload.issuedAt predates it, before allocating an alpha index. - harness-flywheel-generations.ts: the daemon generations loop computed its sequential-evidence testIndex via an unlocked loadAttempts(root).length+1 read before appending. Two overlapping runFlywheelGeneration calls on the same root could be assigned the same test index and spend the same alpha_k twice. The read-index -> build-bundle -> append critical section now runs under the same O_EXCL lock pattern flywheel-transaction.ts already uses. - harness-flywheel.ts: evaluateFlywheelCandidate's sequentialPreflight (backing the `promotable` flag) was computed from an unlocked snapshot taken before the async retrieval/scoring work, so it could go stale by the time promoteFlywheelCandidate allocates the real index under its own lock. Moved to the latest possible read (after receipt registration) to minimize the window; documented as advisory, since promoteFlywheelCandidate remains the sole authority. Added a regression test for the epoch-boundary refusal and a concurrency test proving two racing generations get distinct sequential test indices. Co-Authored-By: RuFlo <ruv@ruv.net> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
fabcc9261a |
fix(hooks): Codex hooks.json schema + PreToolUse verdict compat (#2857)
Codex's plugin hook-manifest loader accepts only `description` and
`hooks` at the top level and hard-rejects the rest of the manifest on
anything else. plugins/ruflo-core/hooks/hooks.json and
plugins/ruflo-cost-tracker/hooks/hooks.json carried `_note` /
`_platform_note` documentation fields, so a fresh Codex install of
ruflo-core@ruflo (still true as of the currently-published 3.32.38 /
ruflo-core 0.2.5) fails to load the plugin at all with:
failed to parse plugin hooks config .../hooks.json:
unknown field `_note`, expected `description` or `hooks`
Fold the doc content into `description` (content preserved, no hook
commands changed) for both marketplace-listed manifests, and add a
strict top-level-key check to
scripts/audit-plugin-hooks-cross-platform.mjs so this class of
regression fails CI going forward. `.claude-plugin/hooks/hooks.json`
is POSIX-only and not marketplace-distributed (not what Codex fetches
for ruflo-core@ruflo); its doc fields are folded too but its
audit-script flags (`_platform`, `_legacy_unaudited_shim`) are kept.
Second, deeper bug once the manifest loads at all: `modify-bash`/
`modify-file` PreToolUse hooks always echo Cursor's
`{"permission":"allow"}` verdict. Codex's own strict output schema
(additionalProperties: false) rejects that shape outright and reports
"hook returned invalid pre-tool-use JSON output" on every single tool
call — verified directly against the real parser
(codex-rs/hooks/src/engine/output_parser.rs,
codex-rs/hooks/src/events/pre_tool_use.rs) and its generated JSON
schema. Empty stdout is the one input Codex's parser treats as
no-opinion/implicit-allow with no error. `isCodexPluginHost()` (added
for #2816, ruflo-core 0.2.5) already detects Codex via its
PLUGIN_ROOT/PLUGIN_DATA env vars and correctly suppresses the verdict
in that case — harden it with a `turn_id`-based fallback (a documented
Codex-only field always present in Codex's PreToolUse input JSON, per
codex-rs/hooks/schema/generated/pre-tool-use.command.input.schema.json)
so detection doesn't depend solely on those env vars being set.
Bump ruflo-core 0.2.5 -> 0.2.6 and ruflo-cost-tracker 0.26.2 -> 0.26.3
so Codex's per-version plugin cache invalidates and re-fetches the
fixed manifest instead of serving a stale cached copy indefinitely.
Fixes #2855, #2856.
|
||
|
|
401e02d511 |
fix: complete reports and consistent initialization for v3.32.37 (#2851)
* fix(metaharness): preserve readiness verdict payloads * test(metaharness): cover blocked genome verdicts * fix(adr): parse bullet metadata and relationships (#2659) * fix(adr): align adr-create with AgentDB schema (#2651) * fix(adr): make index updates idempotent (#2660) * fix(memory): bound session-end graph consolidation (#2628) * fix(memory): align active row visibility (#2652) * fix(memory): honor database path during init * fix(hooks): keep all shim fallback tags aligned * fix(codex): omit unbacked full-template skills * fix(init): generate complete native dual projects * test(memory): isolate path and legacy-row regressions * chore(release): prepare v3.32.37 |
||
|
|
67ff9898c0 |
fix: resolve current runtime and verification defects for v3.32.36 (#2850)
* fix(cli): make runtime status and learning signals truthful * fix(metaharness): accept padded scan severities * fix(codex): make core hooks and status skill native * fix(security): make witness verification hermetic * fix(metrics): expose grounded learning outcomes * fix(hooks): deduplicate project and plugin events * fix(verification): preserve security audit payloads * fix(runtime): stabilize dual memory skills and MCP schemas * fix(cli): harden helper signing and reflexion health * test(release): align helper and container invariants * chore(release): prepare v3.32.36 * test(funnel): align gates with cold-start seed pool * fix(ci): document MCP environment precedence |
||
|
|
9cf769ccce |
feat: ship adaptive swarm and resolve top runtime issues (#2848)
* feat: ship adaptive swarm and issue fixes * fix: register intentional runtime escape hatches |
||
|
|
ddd27576cb |
fix(release): ship Capability Brain as a self-contained 3.32.30 train (#2829)
* fix(release): bundle policy and Codex runtimes * fix(release): verify self-contained three-package archives |
||
|
|
b08246b04a |
chore(release): bump to 3.32.27 (#2823)
* chore(release): bump to 3.32.27 * chore(release): publish policy security dependency * chore(release): sign 3.32.27 helper manifest * ci: harden policy release gates |
||
|
|
0e3d412756 |
feat(flywheel): implement ADR-322 promotion loop (#2817)
Adds verified evaluation receipts, atomic compare-and-swap promotion, bounded Darwin/local proposer integration, CLI/MCP surfaces, ADR specifications, and the 3.32.26 release bump. |
||
|
|
db76d67235 |
fix(metaharness): pin @metaharness/darwin + pin-drift guard (ruflo analog of upstream #142/#149) (#2813)
* fix(metaharness): pin @metaharness/darwin + add pin-drift guard (ruflo analog of #142/#149) Upstream agent-harness-generator #142 (darwin caret-locked to 0.2.x, 3 majors behind) and #149 (META_PROXY_VERSION pinned with no watcher) are both silent pin-drift failures. Ruflo had the same latent gap on its own metaharness deps: - distill-oracle.ts invoked `npx --yes @metaharness/darwin` with NO version pin, so the Tier-1 mechanical oracle floated to whatever npm `latest` was — a breaking darwin release could change eval behavior mid-run. Pinned via a new `MH_DARWIN_PIN = '0.8.0'` constant used by all three npx call sites. - Declared `@metaharness/darwin: ^0.8.0` in optionalDependencies (subprocess- invoked, kept optional per ADR-150/321) so the pin has a single source of truth and installs cache-warm. - Fixed a stale `@metaharness/darwin@~0.3.1` doc comment (darwin is 0.8.0). New guard (the #149 analog ruflo lacked): - scripts/check-metaharness-pins.mjs — diffs each declared range (metaharness, @metaharness/router, @metaharness/darwin) against npm `latest`, plus a lock-step check that MH_DARWIN_PIN satisfies the declared darwin range. Exit 1 on drift; network flakes surface as a warning, never a false positive. - .github/workflows/metaharness-pin-drift.yml — runs the guard weekly + on PRs touching the pins; opens/updates a tracking issue when a pin falls behind, and hard-fails PRs that introduce drift. All pins are current (guard exits 0); router API-surface compat 9/9; CLI builds clean. Lockfile reconciled for the new optionalDep (frozen-lockfile verified). Refs: ruvnet/agent-harness-generator#142, #149; ADR-150, ADR-321. Co-Authored-By: RuFlo <ruv@ruv.net> * chore(release): bump to 3.32.25 (metaharness pin-drift guard) Co-Authored-By: RuFlo <ruv@ruv.net> |
||
|
|
27410d402b |
feat(security): ADR-320 — MCP Composition Inspector v2 + ChannelGuard v2 (#2791)
Follow-up from dream-cycle issue #2783 / PR #2784. Implements the v2 that the already-shipped v1 (commits 381b7ebcc/581cd2bf3) explicitly deferred: SimHash-based cross-tool fragment detection for MCP composition, and a ChannelGuard reusing the real InputValidator instead of a reinvented catalog. 29 new tests, 134/134 existing hooks tests passing, 0 regressions, clippy clean. |
||
|
|
469a901eda |
fix(bridge): wire bridgeRecordFeedback to real intelligence.recordTrajectory (#2786 fix-3)
The prior code called `learningSystem.recordFeedback`/`.record` and
`reasoningBank.recordOutcome`/`.record` — none of those methods exist
on the LocalSonaCoordinator / LocalReasoningBank instances the bridge
actually wires into the registry (see initializeIntelligence in
memory/intelligence.ts). Two silent `catch { /* API mismatch — skip */ }`
blocks swallowed every call. Feedback recording was 100% a no-op for the
CLI's default in-process intelligence path.
Real fix: call `intelligence.recordTrajectory(steps, verdict)` — the
same public API `hooks_post-command` already uses. It initializes
lazily, embeds the step, drives SONA + pattern distillation.
Also fixed the ReasoningBank pattern-store branch to call the ONE
method that actually exists on LocalReasoningBank (`.store(pattern)`)
with the correct StoredPattern shape.
E2E verified in a fresh scratch cwd (v3.32.10 CLI):
- `memory init`
- 3x `hooks post-task --task ... --store-results true`
- `hooks intelligence stats` — Neural Persistence reports 6 trajectories
on disk + 3 pattern entries. Before this fix: 0.
No mocks, no silent catches on the happy path.
Closes: #2786 fix-3 (the last flagged item from the 2026-07-26 tracker
sweep). Was tagged "architectural" because the sweep agent assumed the
target was the AgentDB LearningSystem/ReasoningBank; turned out the
registry wires the local intelligence classes instead, and the intelligence
module already exposes the correct public API. One-file surgical fix.
Co-Authored-By: RuFlo <ruv@ruv.net>
|
||
|
|
9810d8d9c1 |
fix(statusline): real model name from stdin + worktree version resolution
Fixes #2733, #2742. #2733 — hooks.ts's getUserInfo() hardcoded `const modelName = 'Opus 4.6 (1M context)'`, ignoring the actual active model Claude Code passes on stdin entirely. Cosmetically masked in the default render path (the generated .claude/helpers/statusline.cjs already parses stdin correctly and overrides this), but real for direct/manual `hooks statusline` CLI use or any stdin-parse failure in the wrapper. Fixed by mirroring statusline.cjs's own getModelFromStdin() approach inside hooks.ts, so the CLI subcommand is correct standalone. #2742 — getPkgVersion() in the generated statusline.cjs only probed CWD-relative paths (CWD/node_modules/..., CWD/v3/@claude-flow/cli/...). A linked git worktree has no node_modules of its own (worktrees don't get their own `npm install`), so every probe missed and the version silently fell back to the baked-in default from whenever the helper was last generated. Fixed with a pure-fs worktree-root resolver: a linked worktree's `.git` is a plain FILE containing `gitdir: <main>/.git/ worktrees/<name>`; walk up from CWD, parse the pointer, strip the trailing segment to recover the main repo root, and probe its node_modules/v3 paths too. No `git rev-parse` spawn (statusline renders are latency-sensitive). Caught and fixed a real bug in this same resolver during testing: git writes the gitdir pointer with forward slashes even on native Windows, so a path.sep-based (backslash) marker search silently never matched — normalize to forward slashes before searching. `v3/@claude-flow/cli/.claude/helpers/statusline.cjs` is the source of truth per the #2679 redesign (generateStatuslineScript() reads it and substitutes two tokens); regenerated the propagated root-level copy via scripts/regen-statusline-artifact.mjs, which also needed a small cross-platform fix (dynamic import() of a raw Windows path isn't a valid ESM specifier — wrap with pathToFileURL()). Also corrected a stale/misleading comment in message-transport.ts claiming a fresh install "shows the in-code fallback pool until the first refresh lands" — there has never been an in-code fallback pool since ADR-311 ("zero local promo content"); a fresh install's promo row is genuinely empty until the first background refresh lands. Found while investigating a "promo doesn't show on new installs" report; the underlying fail-closed design and SessionStart-triggered refresh mechanism were confirmed working as intended (this machine's own ~/.ruflo state shows a healthy, recently-rotated promo history) — only the comment was wrong, not the behavior. Left alone (separate, unreferenced, ~500-line legacy implementation predating the #2195 delegation rewrite, no source references in its own package): v3/@claude-flow/mcp/.claude/helpers/statusline.cjs. Verified: 21/21 new + existing statusline tests pass, 122/122 funnel tests pass, 10/10 hooks tests pass, clean tsc --noEmit. Manually confirmed #2733 (real stdin model name renders; malformed/empty stdin falls back to "Claude Code", never the old hardcoded string) and #2742 (a real `git worktree add` scenario resolves the main repo's version instead of falling back) end-to-end. |
||
|
|
0a110aee9f |
fix(audit): register #2721's test-only hook env vars as escape hatches
RUFLO_HOOK_CLI_OVERRIDE and RUFLO_HOOK_DEBUG_STDOUT (both added to plugins/ruflo-core/scripts/ruflo-hook.cjs to let test-hooks.mjs point at a local CLI build and observe its output) tripped the env-var-precedence audit's "CLI flag must win" requirement. Same category as the existing RUFLO_HOOK_SKIP_NPX entry: hook scripts have no CLI-flag surface to attach to (invoked by hooks.json, never a user-typed command), and both are test-only — production never sets them. |
||
|
|
b68ad4ccba |
fix(plugins): make ruflo-core/ruflo-cost-tracker hooks Windows-native (#2721)
Both plugins' hooks.json wrapped every command in `/bin/bash -c '...'`,
which fails outright on native Windows (no such path) -- Codex/Claude
Code report "PreToolUse hook (failed) -- exit code 1" on every tool
call. The `_platform: posix` / "ruflo init overrides this on Windows"
claim in both files was never actually true: Claude Code merges
plugin-declared hooks additively with any init-generated
.claude/settings.json, it doesn't replace them, and there's no `ruflo
init` step at all in the reported Codex marketplace install flow.
Fix: every hook command is now a `node -e` bootstrap that resolves
plugins/*/scripts/ruflo-hook.cjs from process.env.CLAUDE_PLUGIN_ROOT
inside Node -- no shell env-var expansion (${VAR} vs %VAR%), so the
exact same command string runs unchanged on Windows/macOS/Linux.
ruflo-core's ruflo-hook.cjs (previously a full port of ruflo-hook.sh
that existed on disk but was never referenced by hooks.json) gained:
- JSON parsing of the hook event from stdin (replaces jq) for
post-command/post-edit, deriving the same CLI flags the bash
version computed
- the PreToolUse permission-allow stdout echo Cursor's stricter
contract requires (previously only the bash wrapper's trailing
printf did this)
- precompact-manual/precompact-auto guidance text (previously plain
bash echoes, no CLI call)
- a real Windows shell-quoting fix: shell:true with an args array
does NOT quote array elements, so "echo hi" silently truncated to
"echo" and a heredoc's `<<` errored as unexpected -- skip the
shell entirely for `node` invocations (never a .cmd shim, so
CreateProcess gets the argv array byte-for-byte)
cost-tracker's existing ruflo-hook.cjs (already correct, just
orphaned) needed no logic changes, only wiring.
Also:
- corrected the false "_platform_note" claims about ruflo init
overriding plugin hooks
- hardened scripts/audit-plugin-hooks-cross-platform.mjs: a
POSIX-exempt hooks.json now must actually reference its sibling
.cjs shim, not just have one sitting on disk unreferenced (which
is exactly the shape cost-tracker shipped in undetected)
- added windows-latest to the plugin-hooks-smoke CI matrix (it was
ubuntu/macos-only because the old bash-based hooks.json couldn't
run on Windows at all) and rewrote test-hooks.mjs to drive hooks.json's
literal command strings via `shell: true` -- exactly how Claude
Code/Codex invoke them -- instead of wrapping everything in an
explicit `bash -c` that could never have caught this bug
- flagged (not fixed) a separate, currently-published, actively
maintained plugin package (.claude-plugin/ + plugin/, the older
"claude-flow" plugin, not listed in the ruflo marketplace) with
the same underlying bug via jq/xargs pipes instead of bash --
explicitly marked _legacy_unaudited_shim so the hardened audit
doesn't silently regress on out-of-scope work
Verified locally on native Windows (this fix's actual target
platform): all 17 ruflo-core hook cases pass, all 3 cost-tracker
cases pass, the existing 12-case smoke-ruflo-hook-cjs.mjs passes
unchanged, both hook-command audits pass clean.
Fixes #2721
|
||
|
|
a4a7d99c22 | test(statusline): document hook-only identity setting | ||
|
|
1fb874005c | test(plugins): preserve standalone MCP catalog assertions | ||
|
|
5e66f065e9 | test(plugins): align namespace and stable hook shims |