DeepSeek-Agent-Harness-2026终极指南-第8章第38节-AgentLoop从零实现-流式Agent:边生成边执行的体验升级
发布时间:2026/10/2 16:56:57来源:尧图网络
DeepSeek Agent Harness 2026终极指南 - 第8章第38节 流式Agent边生成边执行的体验升级第35-37节的Agent Loop是同步的——每次调模型都要等模型生成完才能继续。用户体验很差尤其是长回答要等很久。这节做流式Agent用streamTrue逐token渲染回答同时处理tool_calls的分片重组难题一个工具调用可能拆成多个chunk。从此Agent回答像ChatGPT一样实时流出不再干等。本文导航同步 vs 流式用户体验的天壤之别流式模式下的chunk结构tool_calls分片重组最难的坑流式Agent Loop实现踩坑实录那些流式模式里的坑完整实现streaming_agent.py实测流式回答工具调用小结同步 vs 流式用户体验的天壤之别先看同步模式的问题。假设用户问写一首关于春天的诗模型要生成200个token同步模式发送请求等待10秒模型生成200个token一次性返回200个token用户看到完整回答这10秒里用户盯着空白屏幕不知道模型在干嘛以为卡死了。流式模式发送请求0.5秒后开始收到第一个chunk“春”0.6秒后收到第二个chunk“风”0.7秒后收到第三个chunk“送”…10秒后收到最后一个chunk“。”用户看到文字实时流出流式模式下用户0.5秒就能看到第一个字体验完全不同。流式模式下的chunk结构用streamTrue调模型时返回的不是一个完整的ChatCompletion对象而是一个迭代器每次yield一个ChatCompletionChunkstreamclient.chat.completions.create(modeldeepseek-flash,messages[{role:user,content:你好}],streamTrue,)forchunkinstream:print(chunk)每个chunk的结构{id:chatcmpl-xxx,object:chat.completion.chunk,created:1234567890,model:deepseek-flash,choices:[{index:0,delta:{role:assistant,// 只在第一个chunk出现content:你// 本次chunk的文本片段},finish_reason:null// 只在最后一个chunk有值}]}关键点delta.content是本次chunk的文本片段不是累积的。需要自己拼接。delta.role只在第一个chunk出现后续chunk没有。finish_reason只在最后一个chunk有值如stop其他chunk是null。对于纯文本回答拼接很简单contentforchunkinstream:ifchunk.choices[0].delta.content:contentchunk.choices[0].delta.contentprint(chunk.choices[0].delta.content,end,flushTrue)但如果有tool_calls事情就复杂了。tool_calls分片重组最难的坑当模型决定调用工具时delta里会有tool_calls字段。但tool_calls的分片比content复杂得多// 第1个chunk{delta:{tool_calls:[{index:0,id:call_abc123,function:{name:get_weather,arguments:},type:function}]}}// 第2个chunk{delta:{tool_calls:[{index:0,function:{arguments:{\city\:}}]}}// 第3个chunk{delta:{tool_calls:[{index:0,function:{arguments:\北京\}}}]}}关键设计index字段标识这是第几个tool_call。如果模型同时调两个工具会有index0和index1两个序列。id和name只在第一个chunk出现后续chunk只有arguments片段。arguments是字符串片段需要按index分组拼接。重组算法tool_calls_map{}# {index: {id: ..., name: ..., arguments: }}forchunkinstream:ifchunk.choices[0].delta.tool_calls:fortc_deltainchunk.choices[0].delta.tool_calls:idxtc_delta.indexifidxnotintool_calls_map:# 第一个chunk初始化tool_calls_map[idx]{id:tc_delta.id,name:tc_delta.function.name,arguments:,type:function,}# 拼接argumentsiftc_delta.function.arguments:tool_calls_map[idx][arguments]tc_delta.function.arguments# 转成列表tool_calls[tool_calls_map[idx]foridxinsorted(tool_calls_map.keys())]这个算法的核心是用index作为key把同一个tool_call的所有chunk分到一组第一个chunk提取id和name后续chunk只拼接arguments最后按index排序转成列表流式Agent Loop实现把流式渲染和tool_calls重组整合到Agent Loop里defrun_streaming(user_query:str)-str:流式Agent Loopmessages[{role:system,content:你是一个有用的助手。},{role:user,content:user_query},]toolsregistry.to_openai_tools()foriterationinrange(1,MAX_ITERS1):logger.info(fLoop 第{iteration}轮 ↻)# 流式调用streamclient.chat.completions.create(modelsettings.deepseek_model,messagesmessages,toolstools,streamTrue,)# 拼接content和tool_callscontenttool_calls_map{}finish_reasonNoneforchunkinstream:deltachunk.choices[0].delta finish_reasonchunk.choices[0].finish_reason# 拼接contentifdelta.content:contentdelta.contentprint(delta.content,end,flushTrue)# 拼接tool_callsifdelta.tool_calls:fortc_deltaindelta.tool_calls:idxtc_delta.indexifidxnotintool_calls_map:tool_calls_map[idx]{id:tc_delta.id,name:tc_delta.function.name,arguments:,type:function,}iftc_delta.function.arguments:tool_calls_map[idx][arguments]tc_delta.function.argumentsprint()# 换行# 没有tool_calls → 回答完毕ifnottool_calls_map:returncontent# 有tool_calls → 执行工具tool_calls[tool_calls_map[idx]foridxinsorted(tool_calls_map.keys())]messages.append({role:assistant,content:content,tool_calls:tool_calls,})fortcintool_calls:func_nametc[name]func_argsjson.loads(tc[arguments])logger.info(f → 调用工具:{func_name}({json.dumps(func_args,ensure_asciiFalse)}))resultexecutor.execute(func_name,func_args)logger.info(f ← 工具结果:{result[:80]}{...iflen(result)80else})messages.append({role:tool,tool_call_id:tc[id],content:result,})logger.warning(fAgent Loop 达到最大迭代次数{MAX_ITERS})return[Agent] 抱歉处理超时。关键改动streamTrue开启流式模式遍历chunk拼接content和tool_callsprint(delta.content, end, flushTrue)实时输出最后把拼接好的tool_calls转成列表传给executor.execute()踩坑实录那些流式模式里的坑流式模式有几个容易踩的坑我全踩过了坑1第一个chunk的delta.content可能是空字符串有些模型在第一个chunk只返回{role: assistant}content是空字符串。如果不判断直接拼接会多一个空行。# 错误写法contentdelta.content# 如果delta.content是会多一个空行# 正确写法ifdelta.content:contentdelta.content坑2finish_reason可能为null只有最后一个chunk的finish_reason有值其他chunk是null。如果不判断直接赋值会把null覆盖掉之前的值。# 错误写法finish_reasonchunk.choices[0].finish_reason# 可能被null覆盖# 正确写法ifchunk.choices[0].finish_reason:finish_reasonchunk.choices[0].finish_reason坑3tool_calls的index可能不连续如果模型同时调三个工具index可能是0, 1, 2但也可能是0, 2, 5虽然少见。所以要用index作为key不能用列表下标。# 错误写法tool_calls_list[]fortc_deltaindelta.tool_calls:tool_calls_list.append(...)# 如果index不连续顺序会乱# 正确写法tool_calls_map{}fortc_deltaindelta.tool_calls:idxtc_delta.index tool_calls_map[idx]...坑4流式模式下usage信息可能缺失有些API在流式模式下不返回usageprompt_tokens、completion_tokens只在最后一个chunk返回。如果需要在每次调用后统计token流式模式可能拿不到。解决方案接受这个限制流式模式下不统计token或者在流式结束后再调一次非流式API获取usage浪费一次调用或者用tiktoken自己算不准确但能用我们选择方案1流式模式下不统计token只在非流式模式下统计。完整实现streaming_agent.py把流式Agent Loop整合成完整模块# deep_pilot/streaming_agent.py —— 流式Agent Loop v0.3from__future__importannotationsimportjsonfromtypingimportAnyfromdeep_pilot.clientimportclientfromdeep_pilot.loggerimportget_loggerfromdeep_pilot.tool_registryimportregistryfromdeep_pilot.tool_executorimportget_executorfromdeep_pilot.configimportsettingsimportdeep_pilot.tools# noqa: F401loggerget_logger(__name__)executorget_executor(registry)MAX_ITERS10defrun_streaming(user_query:str)-str: 流式Agent Loop逐token渲染回答实时执行工具。 返回最终的文本回答。 messages[{role:system,content:你是一个有用的助手。当需要查询实时数据时使用提供的工具。},{role:user,content:user_query},]toolsregistry.to_openai_tools()foriterationinrange(1,MAX_ITERS1):logger.info(fLoop 第{iteration}轮 ↻)# 流式调用streamclient.chat.completions.create(modelsettings.deepseek_model,messagesmessages,toolstools,streamTrue,)# 拼接content和tool_callscontenttool_calls_map:dict[int,dict[str,Any]]{}forchunkinstream:deltachunk.choices[0].delta# 拼接contentifdelta.content:contentdelta.contentprint(delta.content,end,flushTrue)# 拼接tool_callsifdelta.tool_calls:fortc_deltaindelta.tool_calls:idxtc_delta.indexifidxnotintool_calls_map:tool_calls_map[idx]{id:tc_delta.id,name:tc_delta.function.name,arguments:,type:function,}iftc_delta.function.arguments:tool_calls_map[idx][arguments]tc_delta.function.argumentsprint()# 换行# 没有tool_calls → 回答完毕ifnottool_calls_map:returncontent# 有tool_calls → 执行工具tool_calls[tool_calls_map[idx]foridxinsorted(tool_calls_map.keys())]messages.append({role:assistant,content:content,tool_calls:tool_calls,})fortcintool_calls:func_nametc[name]func_argsjson.loads(tc[arguments])logger.info(f → 调用工具:{func_name}({json.dumps(func_args,ensure_asciiFalse)}))resultexecutor.execute(func_name,func_args)logger.info(f ← 工具结果:{result[:80]}{...iflen(result)80else})messages.append({role:tool,tool_call_id:tc[id],content:result,})logger.warning(fAgent Loop 达到最大迭代次数{MAX_ITERS})return[Agent] 抱歉处理超时。实测流式回答工具调用uv run python-c from deep_pilot.streaming_agent import run_streaming # 测试1纯文本回答流式渲染 print( 测试1纯文本回答 ) answer run_streaming(写一首关于春天的诗) print(f\n最终答案长度: {len(answer)} 字符) print() # 测试2工具调用流式工具执行 print( 测试2工具调用 ) answer run_streaming(北京今天天气怎么样) print(f\n最终答案: {answer}) 控制台输出 测试1纯文本回答 2026-09-12 19:00:01 | INFO | streaming_agent | Loop 第 1 轮 ↻ 春风轻拂柳丝长 桃李芬芳满院香。 燕子归来寻旧巷 莺歌婉转绕池塘。 最终答案长度: 48 字符 测试2工具调用 2026-09-12 19:00:02 | INFO | streaming_agent | Loop 第 1 轮 ↻ 2026-09-12 19:00:02 | INFO | streaming_agent | → 调用工具: get_weather({city: 北京}) 2026-09-12 19:00:02 | INFO | streaming_agent | ← 工具结果: 晴28°C湿度 45%北风 3 级 2026-09-12 19:00:02 | INFO | streaming_agent | Loop 第 2 轮 ↻ 北京今天天气晴朗气温28°C湿度45%北风3级。适合户外活动。 最终答案: 北京今天天气晴朗气温28°C湿度45%北风3级。适合户外活动。注意测试1的输出是实时流出的不是一次性打印。你在终端里会看到文字一个字一个字地出现体验跟ChatGPT一样。测试2里模型先调工具查天气工具执行完后模型继续生成回答回答也是流式输出的。小结流式模式大幅提升用户体验0.5秒看到第一个字而不是等10秒看完整回答。chunk结构delta.content是文本片段需要拼接delta.tool_calls是工具调用片段需要按index分组拼接。tool_calls分片重组用index作为key第一个chunk提取id和name后续chunk只拼接arguments。踩坑实录空字符串判断、finish_reasonnull判断、index不连续、流式模式下usage缺失。流式Agent Loop遍历chunk → 拼接content和tool_calls → 实时print → 执行工具 → 回填 → 循环。DeepPilot v0.3流式Agent完成——从同步等待到流式渲染用户体验质的飞跃。下节预告流式Agent搞定了但到现在我们还没写过一行测试代码。Agent Loop逻辑复杂多轮循环、工具调用、异常处理手动测试覆盖不全。下一节做单元测试用MockClient替身测试Agent Loop不发真实请求零成本pytest fixture设计循环逻辑边界用例零工具/多工具/异常。从此改代码不怕改坏。如果觉得本文对你有帮助欢迎点赞、收藏、关注三连本系列持续更新中关注不迷路~
网站建设高端定制企业官网