你的 AI Agent 真的在受控运行吗?用 OpenTelemetry 打通 Session 审计日志
发布时间:2026/9/28 18:41:25来源:尧图网络
1. 当 Agent 开始自己动手你还能说清它干了什么吗AI Agent 和普通后端服务最大的区别在于它的行为是非确定的。同一句用户输入模型可能这次调用read读文件下次调用exec跑命令再下次直接发一条 HTTP 请求出去。你没法像审计 REST API 那样靠代码审查预判所有路径因为路径是模型在运行时现编的。这就带来一个很现实的问题当有人问你「你的 Agent 真的在受控运行吗」你拿什么回答谁触发了这次调用、花了多少钱、动了哪些高危工具、行为能不能回放——这四个问题答不上来就谈不上受控。我试过只靠运行时防护工具策略、命令白名单、循环检测来兜底结论是这些机制属于同一信任域内的执行时校验能挡住已知路径但挡不住配置写错、规则漏配、以及上下文压缩把安全指令「挤掉」这类情况。真正能兜底的是一套独立于防护层的可观测体系。这篇就以 OpenClaw 为例把 Session 审计日志、应用日志、OpenTelemetry 遥测三条管道接起来重点落在可复制的配置片段和验证动作上启动 Agent 后触发一次工具调用确认 trace/span 与审计日志同时落库并能按 session 回放完整调用链。适合正在给 Agent 做可观测性落地的同学也适合想搞清楚「审计日志到底该记什么」的运维和安全同学。2. 前置准备TaoToken 与 OpenClaw 环境2.1 为什么这里会提到 TaoTokenOpenClaw 本身是个 Agent 平台它要跑起来得接一个模型服务。TaoToken 在这里的角色是模型接入层你通过它拿到统一的 API Key 和 endpointOpenClaw 的 provider 配置指向它模型调用就能跑通。它不替代 OpenClaw也不替代你的可观测后端只是把「模型从哪来」这件事解决掉。如果你还没配模型先去控制台建一个 Key控制台入口https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentconsoleAPI Key 管理https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentapi-keys拿到 Key 之后OpenClaw 的 provider 配置大致长这样~/.openclaw/openclaw.json片段{ providers: { taotoken: { type: openai-compatible, baseUrl: https://taotoken.net/api, apiKey: sk-你的Key, models: [claude-4-sonnet, gpt-4o] } } }注意 baseUrl 用的是https://taotoken.net/api不带任何查询参数。Key 建议走环境变量注入别硬编码进配置文件。2.2 OpenClaw 侧要开的东西OpenClaw 内置了diagnostics-otel插件这是本篇的核心。先确认插件状态openclaw plugins list openclaw plugins enable diagnostics-otel预期输出里diagnostics-otel的状态是loaded。如果显示disabled说明 enable 没生效检查一下openclaw.json里plugins.allow有没有把它列进去。2.3 可观测后端选型三条管道对后端的要求不一样数据管道生产者传输方式后端要求Session 审计日志SessionManager磁盘 JSONL 采集器支持 JSON 嵌套字段解析应用运行日志tslog Logger磁盘 JSONL 采集器支持按 subsystem 聚合OTEL 遥测diagnostics-otelOTLP/HTTP Protobuf原生支持 OTLP 接入前两条走文件采集LoongCollector、Filebeat、Fluentd 都行第三条直接 OTLP 推送。选后端时优先看它 OTLP 支持是否原生、JSON 嵌套字段查询是否方便这两点决定了后面大盘好不好做。3. 可复制配置三条管道怎么接3.1 Session 审计日志先搞清楚它长什么样Session 日志是审计的核心数据源。每个会话对应一个.jsonl文件每行一个 JSON 对象靠type字段区分条目。一次典型的「读文件」对话会产生这样一串{type:message,id:70f4d0c5,parentId:b5690259,message:{role:user,content:[{type:text,text:帮我读取 /etc/passwd 文件}]}} {type:message,id:3878c644,parentId:70f4d0c5,message:{role:assistant,content:[{type:toolCall,id:call_d46c7e2b,name:read,arguments:{path:/etc/passwd}}],provider:anthropic,model:claude-4-sonnet,usage:{totalTokens:2350},stopReason:toolUse}} {type:message,id:81fd9eca,parentId:3878c644,message:{role:toolResult,toolCallId:call_d46c7e2b,toolName:read,content:[{type:text,text:root:x:0:0:root:/root:/bin/bash\n...}],isError:false}} {type:message,id:a025ab9e,parentId:81fd9eca,message:{role:assistant,content:[{type:text,text:文件内容如下节选...}],usage:{totalTokens:12741,cost:{total:0.0401}},stopReason:stop}}这四行已经能回答谁user让 Agent 做了什么read 读 /etc/passwd、用了哪个模型claude-4-sonnet、花了多少$0.0401、结果如何成功读取。parentId把整条链串起来这就是「按 session 回放」的基础。采集侧配置以 Fluentd 为例LoongCollector 同理source type tail path /root/.openclaw/sessions/*.jsonl pos_file /var/log/td-agent/openclaw-session.pos tag openclaw.session parse type json /parse /source match openclaw.session type http endpoint http://你的日志后端/ingest format type json /format /match关键点type json必须开否则嵌套的message.content会被当成字符串后面没法做字段级查询。3.2 应用日志tslog 的结构化输出应用日志走的是 tslog格式和 Session 日志不同{0:{\subsystem\:\gateway/channels/telegram\},1:webhook processed chatId123456 duration2340ms,_meta:{logLevelName:INFO,date:2026-02-27T10:00:05.123Z,name:openclaw,path:{filePath:src/telegram/webhook.ts,fileLine:142}},time:2026-02-27T10:00:05.123Z}几个关键字段要建索引_meta.logLevelName级别、_meta.path源码定位、数字键0bindings含 subsystem、数字键1消息正文。日志按天滚动文件名openclaw-YYYY-MM-DD.log24 小时自动清理单文件上限 500 MB。采集配置和 Session 日志类似只是 path 换成/root/.openclaw/logs/openclaw-*.logtag 换成openclaw.app。3.3 OTEL 遥测diagnostics-otel 配置这是三条管道里唯一走网络推送的。在~/.openclaw/openclaw.json里加{ plugins: { allow: [diagnostics-otel], entries: { diagnostics-otel: { enabled: true } } }, diagnostics: { enabled: true, otel: { enabled: true, endpoint: YOUR_OTLP_ENDPOINT, protocol: http/protobuf, headers: { Authorization: Bearer YOUR_TOKEN }, serviceName: openclaw-gateway, traces: true, metrics: true, logs: false, sampleRate: 1.0, flushIntervalMs: 60000 } } }几个参数说明一下protocol用http/protobuf别用 gRPC很多后端对 HTTP 支持更稳sampleRate调试期设 1.0生产环境按量调flushIntervalMs默认 60000调试时可以调到 5000 加快验证logs先关掉日志走文件采集更可控。注意endpoint填的是 OTLP 接收地址不是后端首页。填错的话插件会静默失败日志里只有一行 WARN很容易漏。4. 验证触发一次工具调用看数据有没有落库配置写完不算完得验证。验证分三步先确认插件加载再触发一次真实工具调用最后确认 trace 和审计日志同时落库。4.1 确认插件加载openclaw plugins list | grep diagnostics-otel预期输出diagnostics-otel loaded enabled如果是error看~/.openclaw/logs/下当天的日志搜diagnostics-otel关键字通常是 endpoint 或 token 的问题。4.2 触发一次工具调用启动 Gatewayopenclaw gateway start然后发一条会触发工具调用的消息比如让它读一个文件curl -X POST http://localhost:3000/tools/invoke \ -H Authorization: Bearer 你的Gateway Token \ -H Content-Type: application/json \ -d {tool:read,arguments:{path:/etc/hostname}}或者直接在对话界面里发「帮我读一下 /etc/hostname」。关键是这次调用要产生toolCall和toolResult两个条目。4.3 确认 trace/span 与审计日志同时落库先看 Session 日志有没有新文件ls -lt ~/.openclaw/sessions/ | head -5 tail -f ~/.openclaw/sessions/最新文件.jsonl你应该能看到toolCall和toolResult两行parentId串成一条链。再看 OTEL 侧。如果后端支持 trace 查询按serviceNameopenclaw-gateway过滤应该能看到一个 span名字类似openclaw.run带openclaw_model、openclaw_provider等属性。Metrics 侧查一下sum(rate(openclaw_tokens[5m])) by (openclaw_model)如果这条查询有数据说明 OTEL 管道通了。4.4 按 session 回放完整调用链这是验证的最后一环也是最有价值的一环。用 Session 日志里的id和parentId做链式查询type: message and message.role: user | extend content cast(json_extract(message, $.content) as arrayjson) | project content, timestamp, id | unnest | extend content_type json_extract_scalar(content, $.type), content_text json_extract_scalar(content, $.text) | where content_type text拿到 user 消息的id后用parentId 该id查下一条依次往下就能还原整条链user → assistant(toolCall) → toolResult → assistant(stop)。这就是「按 session 回放」的实现方式不需要额外的链路追踪系统Session 日志自带的父子关系就够了。5. 本篇常见错排查5.1 OTEL 插件加载了但没数据最常见的原因是 endpoint 填错。diagnostics-otel对 endpoint 的格式有要求必须是完整的 OTLP 接收路径比如https://your-backend.com/v1/otlp不能只填域名。另外protocol如果填了grpc而后端只支持 HTTP也会静默失败。排查动作把flushIntervalMs调到 5000重启 Gateway看~/.openclaw/logs/里有没有otel export failed之类的 WARN。5.2 Session 日志有 toolCall 但没有 toolResult说明工具执行本身失败了或者执行超时被中断。看应用日志里subsystem: tools-invoke的记录_meta.logLevelName: WARN or _meta.logLevelName: ERROR | project subsystem 0.subsystem, loglevel _meta.logLevelName | where subsystem like %tools-invoke%如果看到EACCES: permission denied是权限问题如果是ENOENT是路径不存在。这两种在审计上都算「越权访问敏感路径」和「配置错误」要分开处理。5.3 应用日志的 subsystem 字段查不出来tslog 把 bindings 放在数字键0里而且值是 JSON 字符串不是对象。查询时要先解析| extend subsystem json_extract_scalar(0, $.subsystem)直接写0.subsystem在某些后端上能work在另一些上会返回 null取决于后端对嵌套字段的解析策略。稳妥做法是先json_extract_scalar再过滤。5.4 Token 消耗指标突然飙升先别急着改配置按这个顺序查OTEL 指标看是哪个 model 涨的 → 应用日志看那个时间窗口有没有 Webhook 重试 → Session 日志按 session 聚合看是哪个会话在烧 token。常见原因是 Prompt 注入导致上下文被恶意填充或者工具调用陷入循环。Session 日志里stopReason: toolUse连续出现多次且toolName相同基本就是循环了。5.5 高危工具在 Gateway HTTP 场景下被调用成功这属于配置绕过要重点排查。OpenClaw 的 Gateway HTTP 默认禁止exec、write、edit、gateway等工具如果 Session 日志里看到这些工具在 HTTP 场景下执行成功说明plugins.allow或工具策略被改过。查法type: message and message.role: assistant and message.stopReason: toolUse | extend content cast(json_extract(message, $.content) as arrayjson) | project content, timestamp | unnest | extend content_type json_extract_scalar(content, $.type), content_name json_extract_scalar(content, $.name) | where content_type toolCall and content_name in (exec,write,edit,gateway)有结果就说明策略没生效得回去看openclaw.json的plugins.allow列表。6. 把三条管道用起来从告警到根因配置和验证都过了之后日常怎么用固定套路是OTEL 指标发现异常 → 应用日志缩小范围 → Session 日志还原行为链。举个例子。OTEL 告警提示openclaw_webhook_error突增先查应用日志_meta.logLevelName: ERROR | extend subsystem json_extract_scalar(0, $.subsystem) | stats cnt count(1) by subsystem | sort cnt desc假设发现集中在gateway/ws再查具体错误_meta.logLevelName: WARN | extend subsystem json_extract_scalar(0, $.subsystem) | where subsystem gateway/ws | project 1, _meta.date看到reasontoken_mismatch且同一remote短时间大量出现基本可以判定是撞库尝试。这时候再去 Session 日志确认有没有会话被建立、有没有工具被调用就能判断影响范围。这套流程的价值在于每一步都有数据支撑不是靠猜。OTEL 告诉你「有异常」应用日志告诉你「哪里异常」Session 日志告诉你「Agent 具体做了什么」。三源联动才能从「有异常」走到「根因是什么」再到「怎么响应」。如果你还没配模型接入可以从模型对话先跑通一条最小链路https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentchat 。长期跑 Agent 编码任务的Coding Plan 会更省心https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentcoding-plan 。接入文档在 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentdoc API Key 在 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentapi-keys 。最后留一个实操建议先把sampleRate设成 1.0、flushIntervalMs设成 5000把三条管道都验证通再逐步调回生产参数。可观测性这东西配置阶段多花十分钟排障阶段能省两小时。
网站建设高端定制企业官网