Files
cklxx 0f0b3ef9dc perf(agent): stop compile=True from recompiling on every new request shape
compile=True wrapped the model in torch.compile(model) with its defaults, so
the first compile was static and the second distinct shape recompiled it as
dynamic, and dynamo's duck sizing then tied dimensions that were equal at
trace time (rows == markers with four questions of four options), so every
later request where they differed recompiled again. On the real checkpoints
the same ten-call sequence built 4 graphs; it now builds 2 (the second is
torch's own specialisation for a single-row batch).

- laya/_compile.py: compile_model() compiles with dynamic=True, and
  independent_dims() sets torch.fx.experimental._config.use_duck_shape to
  False only while a compiled forward runs, restoring the prior value when
  the last concurrent call returns (the setting is read lazily at trace
  time, so it cannot be set once at load and restored).
- Agent: compile=True uses both; mode and device are unchanged (default
  inductor mode, CPU still compiles).
- tests/test_compile.py (CPU, dynamo backend "eager", no weights): one graph
  across shapes, same outputs as eager, setting restored after errors,
  nesting and concurrent threads; added to the CI pytest lanes.
- benchmarks/bench_compile.py: graphs and per-call cost, this tree vs --stock.

Part 1 of splitting #472.

Claude-Session: https://claude.ai/code/session_016MkQeBZCnaqLJTJRfTtVXV
2026-09-27 00:05:19 +08:00
..