ruv
4618c557a0
darwin-core infra: add --only=<dim> flag to benchmark-intelligence.mjs
...
Per-dimension Darwin agents only need one bench's output. Without --only,
the script ran all 6 sub-benches (~5-10 min wall) and three agents in
tick 1+2+3 timed out trying to slice the JSON output for their dim.
Accepts both --only hnsw,sona and --only=hnsw,sona forms (Darwin agents
prefer the = form).
Quick verify: `node scripts/benchmark-intelligence.mjs --only=hnsw --sizes 5000`
returns only the hnsw block in ~5s.
2026-06-27 11:41:06 -04:00
rUv
08b3623759
fix+perf+docs(intelligence): self-learning audit, hardening fixes, honest perf numbers (3.10.7) ( #2228 )
...
* docs(review): empirical capability audit of the intelligence/self-learning system
6 parallel auditors ran real measurements against built dist (not docs).
- Self-learning loop is REAL, measured end-to-end (pattern confidence +
Q-routing change persist cross-process; MoE gating converges; SONA WASM
adapt 0.0042ms; Int8 3.92x; RaBitQ 32x memory — all CONFIRMED).
- Performance multipliers largely UNSUBSTANTIATED: HNSW 150x-12,500x measured
1.48x peak; Flash Attention 2.49-7.47x fabricated via Math.random() at
runtime; embeddings 75x / RaBitQ 2.70x retrieval never benchmarked.
- 1 critical bug: CLI inverts negative reward (route feedback -r -1.0 -> +1.00).
- Real-but-inert: WASM MicroLoRA apply(), MCP consolidation gradient (synthetic),
hooks_intelligence_learn (cosmetic), CLI trajectory-* (no-op stubs), silent
mock-embedding mislabel.
No source modified — audit branch only. Prioritized fix list in the report.
Co-Authored-By: RuFlo <ruv@ruv.net >
* fix+perf(intelligence): harden self-learning system + replace fabricated/unsubstantiated perf numbers with measured ones
Implements the prioritized fix list from docs/reviews/intelligence-system-audit-2026-05-29.md.
FIXES
- CRITICAL (audit #1 , #2222 follow-up): CLI inverted negative reward — route feedback
-r -1.0 / --reward -1.0 parsed as +1.00, so negative feedback REINFORCED the bad
agent. Fixed in src/parser.ts (isFlagValue accepts negative numeric literals). All
three syntaxes now yield -1.0; +6 regression assertions; 52/52 parser tests pass.
- audit #2 : removed fabricated Flash Attention speedup (randomized/midpoint telemetry)
in attention-coordinator (swarm + integration) -> unmeasured sentinel + unverified.
- audit #3 : generateEmbedding returns backend:onnx|mock, surfaced in bridge_status /
import_claude so mock embeddings are never mislabeled as real ONNX.
- audit #4/#5: trajectory-end no longer feeds EWC a synthetic sine gradient;
hooks_intelligence_learn runs a real distill/consolidate cycle. New test added.
OPTIMIZE (HNSW, measured): root cause — HNSW never ran (no storagePath -> native DB
lock held by daemon -> silent catch{} -> brute force). Fixed vector-db.ts (unique
storagePath, hnswConfig m:32/efC:200, visible fallback warning). Same-harness
before->after: N=5000 0.92x->3.2-4.7x, N=20000 0.95x->1.89x; N=1000 below crossover
(reported honestly); reverted m16/efC64 (recall loss).
BENCHMARK: new scripts/benchmark-intelligence.mjs (HNSW N=1k-50k, Int8 3.84x/cos
0.99999, RaBitQ 32x/0.60ms, SONA 0.0043ms, MoE 0.13->0.88, embedding backend).
DOCS: README + 3x CLAUDE.md + READMEs perf tables now show MEASURED values or
unverified/target; HNSW 150x-12,500x & Flash 2.49-7.47x marked NOT reproduced.
Verification: cli build clean (tsc); 161/161 targeted tests pass.
Co-Authored-By: RuFlo <ruv@ruv.net >
* chore(release): bump 3.10.6 → 3.10.7 (intelligence hardening + honest perf docs)
Co-Authored-By: RuFlo <ruv@ruv.net >
2026-05-29 12:13:53 -04:00