Files
multica/playwright.config.ts
6af2268bd9 MUL-7227 test(perf): measure comment typing while agent runs stream (#8276)
* test(perf): measure comment typing while runs stream (MUL-7227)

The regression that made typing a comment take tens of seconds was a summary
animation replacing a DOM node per streamed message, and nothing we had could
see it. The component suite runs in jsdom, which has no style recalculation, no
layout and no main-thread contention, so several thousand passing tests said
nothing about the thing users felt. #8260 pinned the DOM node reuse; this pins
what the user actually waits for.

One browser scenario, on the real page: app shell, shared CSS, the ProseMirror
composer, inline runs, a production Next.js build. Only HTTP and the WebSocket
are synthetic, so no API server, database, daemon, agent or account is
involved. Nothing disables animations or trims the DOM — the cost being
measured lives in exactly that machinery.

Fixed workload, identical for both builds under comparison: 80 comments with
prose and code, three running agents, 1000 seeded transcript messages each, 60
ticks of live messages 100ms apart, and 172 keystrokes at 25ms. Messages are
scheduled from Node rather than the page, so a stalling build cannot quietly
measure less work than the one it is being compared against.

`typing_elapsed_ms` is what the user waits for: first keystroke to the editor
holding the whole string, main-thread queueing included. Style recalculation,
layout, task time and long tasks say where it went. An empty page types very
fast, so a run whose messages never reached the UI, or whose composer did not
end up holding every character, is reported as `invalid` or `timeout` and never
as a duration — including when the scenario blows its budget, which is the case
whose numbers matter most.

Reports rather than gates. `scripts/perf-compare.mjs` builds and measures two
refs in sequence on one machine, each installed against its own lockfile, with
the spec, fixture and browser always taken from the running tree so the product
is the only difference. The workflow is separate and not required: one sample
per ref cannot separate a small regression from machine noise, and the noise
floor has not been calibrated on CI hardware yet.

Verified against the known regression: restoring the pre-#8260 summary
animation moves style recalculation from ~130ms to 1577ms and typing from
~4.9s to 6.2s, against a spread of ~1% across three runs of the fixed build.
Disconnecting the message injection reports `timeout: run 0 end marker`, and
typing fewer characters reports `timeout: editor content` — neither can pass as
fast. #8260's DOM reuse test is untouched and still green.

Tests and run configuration only: no product change.

Co-authored-by: multica-agent <github@multica.ai>

* test(perf): always build for real when comparing two refs (MUL-7227)

The comparison reports a build time, and turbo's cache made it read as ~1.7s
when the real build is ~60s — understating what this costs CI by a factor of
thirty. Worse, a run where one ref hit the cache and the other missed would put
that difference in the timing split as if it belonged to the products.

`--force` on both sides: two real builds, two comparable numbers.

Co-authored-by: multica-agent <github@multica.ai>

* test(perf): report cached builds instead of forcing real ones (MUL-7227)

Forcing a real build on both sides cost minutes on every local run to remove an
ambiguity that a single line of output removes instead. A restored build is
byte-identical to the one that produced it, so the cache cannot move the numbers
this script collects — only the build time it reports, which is now labelled
when it came from the cache.

It buys nothing on CI either: the runner is cold every time, so both builds are
real whether or not they are forced. Locally a repeat comparison drops from
roughly three minutes to under one.

Co-authored-by: multica-agent <github@multica.ai>

* test(perf): run the comparison on demand, and let it exit (MUL-7227)

The comparison never exited on success. The frontend server is spawned
detached, and its teardown hung off `process.on("exit")` — but a live child
keeps the event loop running, so `exit` never fired and the child was never
killed. The report was written and the process then waited on the server it
was supposed to stop. On CI that ran until the job was cancelled at sixty
minutes; locally it left the servers running after being killed. The failure
path only terminated because it happened to call `process.exit(1)`.

Teardown is explicit now. Each side stops its server, waits until the process
has exited and the port has stopped answering, and removes its worktree before
the next side starts — so the second build is also measured on a machine with
nothing of the first one running. The script then exits explicitly, and the
`exit` and signal handlers remain only as a last resort for cancellation.

Checked on every path, with a harness that records the gap between the report
and the exit and looks for anything left behind: success, a scenario failure
with a server running, an unknown ref, and SIGTERM mid-run all exit promptly
with no server, worktree or temp directory remaining. Run through the same
harness, the previous script wrote its report, then was still running 247s
later and left three servers and two worktrees behind.

The workflow is manual only, capped at fifteen minutes. On CI one comparison
took 340s — 292s of it the two production builds, 26s the two measurements —
which is a lot to spend on every frontend PR for a report that gates nothing.
The routine guard stays the component test from #8260, which pins the exact
failure mode and runs with the normal suite; this is for changes that touch
summary animation, long lists, transcript rendering or global styles.

No change to the scenario or its metrics.

Co-authored-by: multica-agent <github@multica.ai>

* test(perf): only trust this run's report, and wait for the whole server (MUL-7227)

Two ways the comparison could report something that was not true.

A report left in the output directory by an earlier run was read back when
this run's scenario failed before writing one, so two failed scenarios showed
the previous run's timings, were marked a usable comparison, and exited 0. And
a scenario that wrote `ok` and then failed was passed through too: the exit
code was recorded as `spec_failed` and then ignored. Each side's report is now
deleted before the scenario runs, the files this script writes are cleared up
front, and a failing scenario process overrides an `ok` report — whatever
failed after the numbers were written is exactly what nobody has looked at.

And a server that stopped answering was taken for one that had stopped. Any
fetch error counted as a closed port, a timeout included, and the server was
dropped from the last-resort cleanup before anything had been confirmed. With
the pnpm wrapper gone but a child still holding the port, the script exited 0
and left two servers running. Teardown now waits for the whole process group to
exit — escalating to SIGKILL if it outlives SIGTERM — keeps the server
registered until then, and treats the port as free only when a connection is
refused, since a hung server still completes the handshake.

`scripts/perf-compare.test.sh` drives the real script with a fake pnpm — no
build, no browser, 9 seconds — through success, a scenario that writes nothing,
a stale report, `ok` followed by failure, a child that ignores SIGTERM, and an
unknown ref. Each of the three new cases fails against the previous script for
the reason above. It runs in `frontend-build` like the other script tests;
the browser comparison itself stays manual.

`PERF_STOP_GRACE_MS` sets the SIGTERM grace (default 10s) so the test can
exercise the SIGKILL path without waiting on it.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-09-11 13:30:21 +08:00

25 lines
715 B
TypeScript

import "./e2e/env";
import { defineConfig } from "@playwright/test";
export default defineConfig({
testDir: "./e2e",
// The performance scenario has its own config: one worker, no retries, and a
// running production build. It must not be swept up by the ordinary suite.
testIgnore: "**/perf/**",
timeout: 60000,
workers: 1,
retries: 0,
use: {
baseURL: process.env.PLAYWRIGHT_BASE_URL ?? process.env.FRONTEND_ORIGIN ?? "http://localhost:3000",
headless: true,
},
projects: [
{
name: "chromium",
use: { browserName: "chromium" },
},
],
// Don't auto-start servers — they must be running already
// This avoids complexity and port conflicts during testing
});