mirror of
https://github.com/multica-ai/multica.git
synced 2026-09-28 13:23:48 +08:00
main
1
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6af2268bd9 |
MUL-7227 test(perf): measure comment typing while agent runs stream (#8276)
* test(perf): measure comment typing while runs stream (MUL-7227) The regression that made typing a comment take tens of seconds was a summary animation replacing a DOM node per streamed message, and nothing we had could see it. The component suite runs in jsdom, which has no style recalculation, no layout and no main-thread contention, so several thousand passing tests said nothing about the thing users felt. #8260 pinned the DOM node reuse; this pins what the user actually waits for. One browser scenario, on the real page: app shell, shared CSS, the ProseMirror composer, inline runs, a production Next.js build. Only HTTP and the WebSocket are synthetic, so no API server, database, daemon, agent or account is involved. Nothing disables animations or trims the DOM — the cost being measured lives in exactly that machinery. Fixed workload, identical for both builds under comparison: 80 comments with prose and code, three running agents, 1000 seeded transcript messages each, 60 ticks of live messages 100ms apart, and 172 keystrokes at 25ms. Messages are scheduled from Node rather than the page, so a stalling build cannot quietly measure less work than the one it is being compared against. `typing_elapsed_ms` is what the user waits for: first keystroke to the editor holding the whole string, main-thread queueing included. Style recalculation, layout, task time and long tasks say where it went. An empty page types very fast, so a run whose messages never reached the UI, or whose composer did not end up holding every character, is reported as `invalid` or `timeout` and never as a duration — including when the scenario blows its budget, which is the case whose numbers matter most. Reports rather than gates. `scripts/perf-compare.mjs` builds and measures two refs in sequence on one machine, each installed against its own lockfile, with the spec, fixture and browser always taken from the running tree so the product is the only difference. The workflow is separate and not required: one sample per ref cannot separate a small regression from machine noise, and the noise floor has not been calibrated on CI hardware yet. Verified against the known regression: restoring the pre-#8260 summary animation moves style recalculation from ~130ms to 1577ms and typing from ~4.9s to 6.2s, against a spread of ~1% across three runs of the fixed build. Disconnecting the message injection reports `timeout: run 0 end marker`, and typing fewer characters reports `timeout: editor content` — neither can pass as fast. #8260's DOM reuse test is untouched and still green. Tests and run configuration only: no product change. Co-authored-by: multica-agent <github@multica.ai> * test(perf): always build for real when comparing two refs (MUL-7227) The comparison reports a build time, and turbo's cache made it read as ~1.7s when the real build is ~60s — understating what this costs CI by a factor of thirty. Worse, a run where one ref hit the cache and the other missed would put that difference in the timing split as if it belonged to the products. `--force` on both sides: two real builds, two comparable numbers. Co-authored-by: multica-agent <github@multica.ai> * test(perf): report cached builds instead of forcing real ones (MUL-7227) Forcing a real build on both sides cost minutes on every local run to remove an ambiguity that a single line of output removes instead. A restored build is byte-identical to the one that produced it, so the cache cannot move the numbers this script collects — only the build time it reports, which is now labelled when it came from the cache. It buys nothing on CI either: the runner is cold every time, so both builds are real whether or not they are forced. Locally a repeat comparison drops from roughly three minutes to under one. Co-authored-by: multica-agent <github@multica.ai> * test(perf): run the comparison on demand, and let it exit (MUL-7227) The comparison never exited on success. The frontend server is spawned detached, and its teardown hung off `process.on("exit")` — but a live child keeps the event loop running, so `exit` never fired and the child was never killed. The report was written and the process then waited on the server it was supposed to stop. On CI that ran until the job was cancelled at sixty minutes; locally it left the servers running after being killed. The failure path only terminated because it happened to call `process.exit(1)`. Teardown is explicit now. Each side stops its server, waits until the process has exited and the port has stopped answering, and removes its worktree before the next side starts — so the second build is also measured on a machine with nothing of the first one running. The script then exits explicitly, and the `exit` and signal handlers remain only as a last resort for cancellation. Checked on every path, with a harness that records the gap between the report and the exit and looks for anything left behind: success, a scenario failure with a server running, an unknown ref, and SIGTERM mid-run all exit promptly with no server, worktree or temp directory remaining. Run through the same harness, the previous script wrote its report, then was still running 247s later and left three servers and two worktrees behind. The workflow is manual only, capped at fifteen minutes. On CI one comparison took 340s — 292s of it the two production builds, 26s the two measurements — which is a lot to spend on every frontend PR for a report that gates nothing. The routine guard stays the component test from #8260, which pins the exact failure mode and runs with the normal suite; this is for changes that touch summary animation, long lists, transcript rendering or global styles. No change to the scenario or its metrics. Co-authored-by: multica-agent <github@multica.ai> * test(perf): only trust this run's report, and wait for the whole server (MUL-7227) Two ways the comparison could report something that was not true. A report left in the output directory by an earlier run was read back when this run's scenario failed before writing one, so two failed scenarios showed the previous run's timings, were marked a usable comparison, and exited 0. And a scenario that wrote `ok` and then failed was passed through too: the exit code was recorded as `spec_failed` and then ignored. Each side's report is now deleted before the scenario runs, the files this script writes are cleared up front, and a failing scenario process overrides an `ok` report — whatever failed after the numbers were written is exactly what nobody has looked at. And a server that stopped answering was taken for one that had stopped. Any fetch error counted as a closed port, a timeout included, and the server was dropped from the last-resort cleanup before anything had been confirmed. With the pnpm wrapper gone but a child still holding the port, the script exited 0 and left two servers running. Teardown now waits for the whole process group to exit — escalating to SIGKILL if it outlives SIGTERM — keeps the server registered until then, and treats the port as free only when a connection is refused, since a hung server still completes the handshake. `scripts/perf-compare.test.sh` drives the real script with a fake pnpm — no build, no browser, 9 seconds — through success, a scenario that writes nothing, a stale report, `ok` followed by failure, a child that ignores SIGTERM, and an unknown ref. Each of the three new cases fails against the previous script for the reason above. It runs in `frontend-build` like the other script tests; the browser comparison itself stays manual. `PERF_STOP_GRACE_MS` sets the SIGTERM grace (default 10s) so the test can exercise the SIGKILL path without waiting on it. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |