Laya 预测钩子生命周期全解:Agent 与 Router 每次调用的精确事件顺序
发布时间:2026/9/30 2:25:22来源:尧图网络
人工智能NLP强化学习【免费下载链接】layaNon-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100 languages, with a router that picks the right checkpoint per request.项目地址https://gitcode.com/gh_mirrors/lay/laya点击查看免费下载本指南以 Laya 开源仓库中的 docs/hooks/lifecycle.md 为骨架逐事件拆解Agent.predict_batch、system_one/predict、Router.predict、Router.predict_batch以及模型加载/淘汰on_load/on_evict的完整调用顺序并给出缓存命中ctx.skip、空输入、钩子排序规则与并发语义。读完本文你将掌握 Laya 每次决策请求中「哪些钩子在什么时机运行、能否改写什么、失败时如何收尾」并能基于生命周期写出审计、脱敏、缓存、指标上报等生产级钩子。配套的钩子系统总览见 docs/hooks/index.md事件失败策略见 docs/hooks/errors.md写法规范见 docs/hooks/patterns.md。1. 生命周期总览从一次入口调用看事件流水线Laya 的钩子系统围绕一个核心对象——PredictContext——展开。每次调用无论走Agent还是Router都会创建一个独立的PredictContext同一个ctx对象贯穿start、error、end三个阶段因此run_id可以把三个阶段关联起来end 钩子也能读到ctx.error。laya/hooks.py 中定义了完整的生命周期方法集合HOOK_EVENTSon_predict_start推理开始前可改写ctx.states/ctx.questions、设置ctx.max_len/ctx.head_max_len、调用ctx.skip(results)短路推理或直接抛异常中止on_predict_end推理结束后无论成败都会运行可改写ctx.resultson_error失败路径上运行能读到ctx.erroron_route仅 Router路由决策产生后运行可替换ctx.decisionon_load/on_evict仅 Router模型生命周期事件在 Router 内部锁释放之后触发。PredictContext携带states、questions、results、decision、model、agent、router、max_len、head_max_len、usage、elapsed_ms、error以及自动生成的run_id实现见 laya/hooks.py。它是一个eqFalse的 dataclass两个 context 永不相等、按身份可哈希钩子可以把它放进集合而不会误比较内部状态。一句话概括设计意图钩子既可以被安装installed也可以按调用传入per-call两者按固定顺序合并start 钩子负责「塑造」调用end 钩子负责「观察/收尾」ctx.skip提供免推理的缓存通道。2. Agent.predict_batch单实现入口predict_batch是整个 Agent 层的唯一实现system_one与predict都只是带一个 state 的薄封装。其生命周期如下与 docs/hooks/lifecycle.md 的流程图一致实现在 laya/agent.pypredict_batch(states, questions, batch_size..., hooks..., ...) │ ├─ active installed hooks per-call hooks (installed first) ├─ ctx PredictContext(states, questions, modelself.model_id, agentself) │ ├─ try: │ │ │ ├─ on_predict_start ─────────────────────────────┐ │ │ a hook may: │ │ │ • rewrite ctx.states / ctx.questions │ │ │ • set ctx.max_len / ctx.head_max_len │ │ │ • ctx.skip(results) ─────────────┐ │ │ │ • raise (aborts; see errors) │ │ │ │ │ │ │ ├─ if ctx.results is not None: ◄─────────┘ │ cache hit │ │ skip tokenization and forward │ │ ├─ else: │ │ │ validate states is a list │ │ │ for each batch chunk: │ │ │ _encode_state ─► collate ─► _forward │ │ │ _decode_answers │ │ │ ctx.results [...] │ │ │ │ │ └─ (any failure here) ──► except BaseException: │ │ ctx.error exc │ │ on_error │ │ re-raise │ │ │ │ finally: │ │ ctx.elapsed_ms now - ctx.started_at │ │ if ctx.results: ctx.usage aggregate_usage(...)│ │ on_predict_end ─────────────────────────────────┘ │ └─ return ctx.results几个关键实现细节值得注意钩子合并active compose_hooks(self.hooks, hooks, on_predict_start, on_predict_end)其中实例上安装的钩子在前、按调用传入的钩子在后laya/agent.py。compose_hooks还会把进程级默认钩子set_default_hooks注册的排在最前面laya/hooks.py。start 钩子的改写直接生效dispatch之后代码立即读取ctx.states, ctx.questions用于后续编码与推理。测试 tests/test_hooks.py 验证了「start 钩子把 state 改成rewritten实际被编码的就是rewritten」。cache hit 短路如果 start 钩子调用了ctx.skip(results)ctx.results非空tokenization 与 forward 全部跳过但finally中的elapsed_ms、usage聚合与on_predict_end依然执行见 laya/agent.py。失败的统一收尾except BaseException会把异常写入ctx.error、运行on_error后重新抛出finally中先补上耗时与用量再运行on_predict_end——即使推理失败end 钩子也会看到ctx.error。这正是try / except / finally三明治结构docs/hooks/errors.md 中的失败矩阵基于此展开。start 钩子抛错同样触发收尾测试 tests/test_hooks.py 验证 start 钩子抛ValueError时on_error照常触发一次、end 钩子照常运行且能读到该错误。3. Agent.system_one / predict单状态薄封装system_one即predict的别名见 laya/agent.py实现为system_one(state, questions, hooks..., ...) └─ predict_batch([state], questions, hooks..., ...)[0]因此system_one继承全部钩子与同一套生命周期只是ctx.states [state]。也就是说你在system_one上看到的所有钩子行为改写、跳过、错误收尾与predict_batch完全一致无需额外学习两套规则。测试 tests/test_hooks.py 专门验证了system_one会触发安装的钩子。4. Router.predict路由 推理的完整流水线Router.predict是生命周期最完整的入口它把路由、模型加载、推理三层事件串联起来laya/router.pyRouter.predict(state, questions, model..., hooks..., on_predict_start..., on_predict_end...) │ ├─ active installed hooks per-call hooks │ ├─ route(state, questions, ..., hooksper-call, hooks_raise...) │ │ │ ├─ _route(...) detect script / language / workflow │ ├─ on_route ──► ctx.decision a hook may replace the decision │ └─ return ctx.decision │ ├─ load(decision[model]) │ │ │ ├─ already resident? ──► return it │ ├─ else build Agent(...) ──► on_load (after the Router lock is released) │ └─ evict LRU checkpoints ──► on_evict (after the Router lock is released) │ ├─ ctx PredictContext(states[state], questions, decision, modeldecision.model, │ agentagent, routerself) ├─ try: │ ├─ on_predict_start │ ├─ if ctx.results is None: │ │ result agent.system_one(ctx.states[0], ctx.questions) │ │ └─ the Agents own hooks run here (start / forward / end) │ │ result[routing] decision │ │ ctx.results [result] │ └─ else: │ for each cached result: result.setdefault(routing, decision) │ └─ (any failure) ──► except: on_error, re-raise │ └─ finally: elapsed_ms, usage, on_predict_end │ └─ return ctx.results[0]生命周期文档反复强调的三个关键点在源码中都能印证on_route先于模型加载。route()内部的_route只做脚本/语言/工作流检测laya/router.py不触碰任何模型权重钩子可以在此替换ctx.decision例如把某类流量固定到指定 checkpoint从而避免加载另一个模型。测试 tests/test_hooks.py 演示了on_route把english替换为multilingual的过程。Router 级 predict 钩子包裹整个调用但不会转发进 Agent。Router 调用的是agent.system_one(...)内部不传 Router 的钩子被附加的 Agent 如果自己装有钩子会按自身生命周期另行运行——这是预期行为而非缺陷。Router 在调用 Agent 前会临时关闭进程级默认钩子_SKIP_DEFAULTS.set(True)见 laya/router.py避免进程级钩子在同一请求里被 Router 与 Agent 各触发一遍。Router 级ctx.skip()仍然补齐routing键。缓存命中时返回负载通过result.setdefault(routing, dict(decision))补上路由信息保证返回值形状稳定laya/router.py。此外Router.predict会把路由检测到的语言作为lang转发给 Agenteffective_lang使 Agent 的按语言温度配置生效且对不接受lang参数的 Agent 类对象有TypeError容错回退laya/router.py。5. Router.predict_batch批量的逐请求生命周期Router.predict_batch承诺「每个结果等价于对该请求单独调用predict的返回值」因此 Router 级 predict 钩子在这里也是逐请求触发的每个请求拥有独立的PredictContext、run_id和elapsed_mslaya/router.pyRouter.predict_batch(requests, batch_size...) │ ├─ route_batch(requests) ──► on_route, once per request (no checkpoint loaded yet) │ └─ for each checkpoint, in order of first appearance: │ ├─ load(checkpoint) ──► on_load / on_evict ├─ for each request of this checkpoint, in input order: │ ctx PredictContext(states[state], questions, decision, model, agent, router) │ on_predict_start a hook may redact, rewrite, set a token budget or skip ├─ group the requests left to infer by (questions, ctx.max_len, ctx.head_max_len) │ agent.predict_batch(states, questions, ...) ──► one shared forward pass per group │ result[routing] decision; ctx.results [result] ├─ (any failure) ──► on_error for every started request without a result, │ on_predict_end for every started request, re-raise └─ on_predict_end, once per request of this checkpoint, in input order四个必须理解的批量语义均有源码与测试佐证同组请求的 start 全部先于任何 end因为同一 checkpoint 的请求共享 forward pass缓存类钩子若在on_predict_end填充缓存就无法服务同一组内的重复 state——它只能在跨调用时命中文档原意测试 tests/test_hooks.py 展示了每个请求严格的 start→end 配对与事件顺序。start 钩子的改写只影响自己的请求请求在 start 钩子运行之后才按(questions, ctx.max_len, ctx.head_max_len)分组因此改写ctx.states、ctx.questions或 token 预算不会波及其他请求但若原地修改一个共享的 questions 字典则会影响所有共享该字典的请求以及调用方——这与predict的行为一致。测试 tests/test_hooks.py 验证了「只有long请求拿到max_len1024」与「改写后的 questions 单独成组」。每个已 start 的请求恰好触发一次on_predict_end即使前一个请求的 end 钩子抛错所有 end 运行完后第一个错误才被抛出_end_contexts的实现laya/router.py。批量失败时没有结果的已 start 请求以on_error报告失败ctx.error为导致批量失败的异常因为调用方拿不到它们的结果已经完成的 checkpoint 组的请求则正常结束——等价于逐个调用[router.predict(...) for ...]的效果。测试 tests/test_hooks.py 精确验证了这一矩阵含失败组内缓存命中请求不报错、保留结果的场景。分组共享 forward pass 带来的收益在 laya/router.py 中可以看到同一 checkpoint 下按「顺序敏感的 question schema token 预算 lang」再细分同组内一个 forward pass 批量回答多个 state。6. 模型生命周期on_load 与 on_evicton_load在 checkpoint 被构建时触发on_evict在 checkpoint 被释放时触发。两者都在Router 内部锁释放之后运行因此钩子可以安全地回调 Routerlaya/router.pyload(multilingual) │ ├─ [lock] │ build Agent(...) (seconds: download weights) │ register in _agents / _order │ evict LRU if over max_loaded ──► evicted [english] ├─ [unlock] ├─ on_evict(english) └─ on_load(multilingual) unload(english) ├─ [lock] remove from _agents / _order ├─ [unlock] └─ on_evict(english)load()在锁内构建 Agent、登记_agents/_orderLRU 顺序、必要时淘汰超限 checkpoint随后在锁外依次派发on_evict被淘汰者与on_load新构建者。on_evict派发在on_load之前——这是排序规则第 5 条的来源。源码层面_dispatch_lifecycle为每个被淘汰名字单独构造一个PredictContext(states[], questions{}, modelname, routerself)laya/router.py。另一个细节attach(name, agent)注册一个已构建的 Agent 不会触发on_load因为没有构建任何 checkpointlaya/router.py。这在测试 tests/test_hooks.py 中有专门验证max_loaded1下依次load(english)、load(multilingual)产生loads[english,multilingual]、evicts[english]。7. 缓存与 skip免推理的命中路径生命周期文档给出的缓存语义非常清晰on_predict_start ├─ cache hit? ctx.skip([cached_result]) │ └─ forward pass skipped │ └─ on_predict_end still runs │ └─ Router adds routing if missing └─ cache miss? nothing └─ forward pass runs └─ on_predict_end can store the result仓库里有一个可直接运行的完整示例 examples/hooks/cache.pycache_read在 start 阶段按(state, questions)的 SHA-256 摘要查缓存并ctx.skip([hit])cache_write在 end 阶段回填缓存。测试 tests/test_hooks.py 验证了三条不变量skip 返回缓存结果、_forward_calls为空没有 forward pass、end 钩子照常执行。实操要点并发服务时请为缓存加锁见 docs/hooks/patterns.md 的缓存一节Router 场景下ctx.skip的负载会自动补上routing键返回形状与正常推理结果完全一致由于同组请求共享 forward pass批量路径下缓存回填要到下一批调用才能命中见第 5 节。8. 空输入钩子照常触发审计不丢调用predict_batch([])、predict_batch(states, {})空 questions与system_one(state, {})都不会做 tokenization 或 forward pass但on_predict_start与on_predict_end仍然运行因此审计钩子能看到每一次调用。生命周期文档给出的ctx.results终值如下输入结束时ctx.resultspredict_batch([])[]predict_batch(states, {})无 questions每个 state 一个空答案负载system_one(state, {})单个空答案负载源码印证states为空时直接置ctx.results []questions为空时生成{answers: {}, usage: {input_tokens: 0, output_tokens: 0}}的占位结果laya/agent.py。测试 tests/test_hooks.py 验证了空输入下 start 看到resultsNone、end 看到空结果列表且无任何编码与推理调用。9. 排序规则谁先谁后一锤定音生命周期文档给出了五条硬性排序规则安装的钩子永远先于 per-call 钩子同一列表内钩子按列表顺序运行同一事件所有实现它的钩子按顺序全部跑完才进入下一事件失败路径上on_error先于on_predict_end一次既淘汰又构建的loadon_evict先于on_load。示例installed: [A, B] per-call: [C] on_predict_start: A, B, C on_predict_end: A, B, C排序在源码中有两处保证一是compose_hooks的拼接顺序默认 → 安装 → per-call见 laya/hooks.py二是dispatch按顺序对每个 hook 调用事件方法跳过未实现的钩子laya/hooks.py。normalise_hooks还会在构造期校验钩子拒绝类而非实例、拒绝没有任何生命周期方法的对象、拒绝非可调用的事件laya/hooks.py。测试覆盖了核心排序安装钩子与 per-call 钩子顺序为[installed, percall]同一列表内按[1, 2]顺序执行进程级默认钩子跑在安装钩子之前tests/test_hooks.py。10. 并发语义Agent 与 Router 的线程安全Agent与Router都支持多线程并发调用。每次调用创建独立的PredictContextcontext 永远不会跨请求泄漏唯一共享的状态是钩子对象本身。因此一个不是线程安全的钩子要么自己加锁保护内部状态要么以hooks_concurrentFalse安装hooks_concurrentTrue默认时多线程下钩子并行运行hooks_concurrentFalse时通过一个RLock可重入锁让同一事件每次只有一个钩子在跑。hooks_concurrentTrue (default) hooks_concurrentFalse thread 1 ─┐ thread 1 ─┐ thread 2 ─┼─ hooks run in parallel thread 2 ─┼─ one hook at a time thread 3 ─┘ thread 3 ─┘ (RLock)两点容易混淆的语义需要强调hooks_concurrentFalse序列化的是「每次钩子调用」而不是「整个调用」两次调用仍可在事件之间交错执行因为用的是可重入锁钩子可以回调同一个Agent/Router而不会死锁laya/agent.py、laya/router.py。测试 tests/test_hooks.py 验证了默认并发下 5 个线程产生 5 个独立 contextlen(set(contexts)) 5以及hooks_concurrentFalse下max_active 1钩子从未并行。需要留意的是钩子默认同步执行且跑在调用线程上laya.serve使用单一推理 worker阻塞式钩子网络等待、input()会拖慢整个服务——这也是 docs/hooks/patterns.md 把「Blocking work」列为头号反模式的原因。11. 钩子生命周期在 ONNXAgent 与 serve 中的延伸同一个生命周期约定也适用于轻量运行时ONNXAgent.system_one使用相同的compose_hooksdispatch结构同样支持 start 改写与ctx.skip短路见 laya/onnx_agent.py 及测试 tests/test_hooks.py。而laya.serve与 MCP server 内部调用的是Router.predict因此 Router 级钩子含on_route、on_load、on_evict在 HTTP/服务场景下自动生效见 docs/hooks/index.md 的 Scope 说明。12. 实战组合示例把生命周期用完整结合 docs/hooks/patterns.md 与 examples/hooks/ 中的可运行示例一个「脱敏 缓存 审计」的组合可以这样安装到 Router 上import hashlib, json, re from laya import Router EMAIL re.compile(r\b[\w.-][\w-]\.[\w.-]\b) CACHE {} def cache_key(state, questions): return hashlib.sha256( json.dumps([state, questions], sort_keysTrue, defaultstr).encode() ).hexdigest() def redact(ctx): # on_predict_start推理前脱敏 ctx.states [EMAIL.sub([email], s) if isinstance(s, str) else s for s in ctx.states] def cache_read(ctx): # on_predict_start命中即 skip hit CACHE.get(cache_key(ctx.states[0], ctx.questions)) if hit is not None: ctx.skip([hit]) def audit(ctx): # on_predict_end无论成败都记录 if ctx.results is None: print(failed, ctx.run_id, ctx.error) else: print(ctx.run_id, ctx.model, ctx.results[0].get(routing, {}).get(model), round(ctx.elapsed_ms or 0.0, 3)) router Router( on_predict_start[redact, cache_read], # 按列表顺序先脱敏后查缓存 on_predict_endaudit, hooks_raiseFalse, # 审计失败不拖垮请求 )这个例子演示了本篇文章的全部核心start 阶段按安装顺序塑造调用脱敏 → 缓存查询 → skipend 阶段观察并记录结果失败路径由try/except/finally保证on_error与on_predict_end依然运行。若你还需要按事件拆分更细的失败处理策略请继续阅读 docs/hooks/errors.md若想深入每个钩子事件的参数与默认值见 docs/hooks/api.md。赞分享人工智能NLP强化学习【免费下载链接】layaNon-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100 languages, with a router that picks the right checkpoint per request.项目地址https://gitcode.com/gh_mirrors/lay/laya点击查看免费下载相关推荐mini-vue组件生命周期调试钩子执行顺序mini vue组件生命周期调试钩子执行顺序 你是否在开发组件时遇到过钩子函数执行顺序混乱的问题是否想知道为什么数据更新后页面没有立即刷新本文将带你深入理前端教程Vue Native中的生命周期钩子使用场景与执行顺序Vue Native中的生命周期钩子使用场景与执行顺序 你是否曾在开发Vue Native应用时遇到数据初始化时机不当导致的bug或者因组件销毁时未清理资源移动开发跨平台前端NoneBot2 钩子函数Hook完全指南生命周期钩子与事件处理钩子详解NoneBot2 钩子函数Hook完全指南生命周期钩子与事件处理钩子详解 钩子编程hooking也称作“挂钩”是计算机程序设计术语指通过拦截软件后端即时通讯创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
网站建设高端定制企业官网