mirror of
https://github.com/NandhaKishorM/laya.git
synced 2026-09-28 16:02:56 +08:00
compile=True wrapped the model in torch.compile(model) with its defaults, so the first compile was static and the second distinct shape recompiled it as dynamic, and dynamo's duck sizing then tied dimensions that were equal at trace time (rows == markers with four questions of four options), so every later request where they differed recompiled again. On the real checkpoints the same ten-call sequence built 4 graphs; it now builds 2 (the second is torch's own specialisation for a single-row batch). - laya/_compile.py: compile_model() compiles with dynamic=True, and independent_dims() sets torch.fx.experimental._config.use_duck_shape to False only while a compiled forward runs, restoring the prior value when the last concurrent call returns (the setting is read lazily at trace time, so it cannot be set once at load and restored). - Agent: compile=True uses both; mode and device are unchanged (default inductor mode, CPU still compiles). - tests/test_compile.py (CPU, dynamo backend "eager", no weights): one graph across shapes, same outputs as eager, setting restored after errors, nesting and concurrent threads; added to the CI pytest lanes. - benchmarks/bench_compile.py: graphs and per-call cost, this tree vs --stock. Part 1 of splitting #472. Claude-Session: https://claude.ai/code/session_016MkQeBZCnaqLJTJRfTtVXV