mirror of
https://github.com/NandhaKishorM/laya.git
synced 2026-09-28 07:52:57 +08:00
* feat(hooks): opt-in prediction hooks to observe or shape every decision Add on_predict_start/on_predict_end and a Hook protocol to Agent, Router and ONNXAgent, plus on_route/on_load/on_evict on the Router. Hooks receive a mutable PredictContext: a start hook can redact or rewrite the state/questions, or call ctx.skip(...) to serve a cached result and skip the forward pass; an end hook can rewrite the results. Unset hooks are a no-op, so the default path is unchanged. Adds tests/test_hooks.py (50 checks, no weights), examples/hooks/ and docs/hooks.md. * fix(hooks): run start/error/end on start-hook failure; expose model id; per-call on_route Review pass: - A failing on_predict_start hook now still runs on_error and on_predict_end, instead of aborting before the try/finally. - Agent and ONNXAgent populate ctx.model from the checkpoint id, matching the documented context. - Router per-call hooks= now apply to on_route too (and route() takes hooks and hooks_raise), so a per-call hook covers the whole call. - normalise_hooks rejects a class instead of an instance with a clear error. Adds 9 checks to tests/test_hooks.py (59 total). * test(hooks): regression + API-stability tests; fix router cache-hit routing - tests/test_hooks_api.py pins the hook surface (parameter names/defaults, PredictContext fields, lifecycle events, exports, class defaults) so an accidental API change fails CI. - Regression checks for the review-pass fixes, plus hooks_concurrent storage. - Router.predict now adds routing on a cache-hit skip, keeping its documented return contract. - PredictContext uses identity equality/hash so a hook can store it in a set. - normalise_hooks rejects non-callable lifecycle attributes. * feat(hooks): add run_id for tracing; document real-world coverage - PredictContext.run_id is a unique id shared by every hook of one call, so a tracer can correlate start/end/error spans without its own bookkeeping. - docs/hooks.md: prior-art mapping (browser-use on_step_*, OpenAI Agents SDK RunHooks/AgentHooks), a tracing example, and a note that hooks are synchronous and must not block (laya.serve runs inference on a single worker). - tests: run_id regression + API field guard; prod scenario correlating spans. * docs(hooks): expand hooks.md into a docs/hooks/ folder Replace the single hooks.md with a seven-page reference: overview, full API, lifecycle flowcharts, error handling, patterns and anti-patterns, examples and tracing. Every public hook API is covered, links and anchors are checked by test_packaging.py, and the README/examples links point at docs/hooks/index.md. * feat(hooks): runtime registration, scoped install, per-call token budget - HookRegistry mixin: add_hook / remove_hook / hooks_installed on Agent, Router and ONNXAgent. Thread-safe; a call reads a snapshot, so mutation never disturbs a call in flight. - PredictContext.max_len / head_max_len: a start hook, or a per-call max_len= / head_max_len= argument, shapes the token budget for one call without touching shared agent config. Honored by Agent and ONNXAgent, propagated from Router. - Tests: 86 hook checks, 101 API-stability checks; docs updated (API, lifecycle, patterns, examples). --------- Co-authored-by: Nandakishor <48623612+NandhaKishorM@users.noreply.github.com>
Prediction hook examples
Small, runnable hooks. See docs/hooks/index.md for the full reference.
audit.py: log every decision and optionally ship it to an external service.redact.py: strip emails/phones from the state before inference.cache.py: cache decisions and skip the forward pass on a hit.otel.py: counters and a latency histogram per decision.
Each example loads a real checkpoint, so the first run downloads it.