transformers-js - PIPELINE_OPTIONS
发布时间:2026/10/2 13:45:05来源:尧图网络
管道选项参考本指南介绍如何通过pipeline()函数中的PretrainedModelOptions参数配置模型加载和推理。目录概述基本选项模型加载选项设备与性能选项常见配置模式概述pipeline()函数接受三个参数import{pipeline}fromhuggingface/transformers;constpipeawaitpipeline(task-name,// 1. 任务类型例如sentiment-analysismodel-id,// 2. 模型标识符可选为 null 时使用默认模型options// 3. PretrainedModelOptions可选);第三个参数options允许您配置模型的加载和执行方式。可用选项interfacePretrainedModelOptions{// 进度跟踪progress_callback?:(info:ProgressInfo)void;// 模型配置config?:PretrainedConfig;// 缓存与加载cache_dir?:string;local_files_only?:boolean;revision?:string;// 模型特定设置subfolder?:string;model_file_name?:string;// 设备与性能device?:DeviceType|Recordstring,DeviceType;dtype?:DataType|Recordstring,DataType;// 外部数据格式大型模型use_external_data_format?:boolean|number|Recordstring,boolean|number;// ONNX Runtime 设置session_options?:InferenceSession.SessionOptions;}基本选项进度回调跟踪模型下载和加载进度。推荐使用progress_total获取端到端进度可选地使用progress获取逐文件详情。constfileProgress{};constpipeawaitpipeline(sentiment-analysis,null,{progress_callback:(info){// 推荐端到端加载进度if(info.statusprogress_total){console.log(Total:${info.progress.toFixed(1)}%);return;}// 可选逐文件进度if(info.statusprogress){fileProgress[info.file]info.progress;console.log(${info.file}:${info.progress.toFixed(1)}%);}if(info.statusdone){console.log(✓${info.file}complete);}}});进度信息类型typeProgressInfo{status:initiate|download|progress|progress_total|done|ready;name:string;// 模型 ID 或路径file?:string;// 正在处理的文件逐文件事件progress?:number;// 百分比0-100用于 progress 和 progress_totalloaded?:number;// 已下载字节数仅用于 progress 状态total?:number;// 总字节数仅用于 progress 状态};示例带多文件下载的浏览器加载 UIconststatusDivdocument.getElementById(status);constprogressContainerdocument.getElementById(progress-container);constfileProgressBars{};constpipeawaitpipeline(image-classification,null,{progress_callback:(info){if(info.statusprogress_total){statusDiv.textContentTotal:${info.progress.toFixed(1)}%;return;}if(info.statusprogress){// 为每个文件创建进度条如果不存在if(!fileProgressBars[info.file]){constfileDivdocument.createElement(div);fileDiv.innerHTMLdiv classfile-name${info.file}/div div classprogress-bar div classprogress-fill stylewidth: 0%/div /div;progressContainer.appendChild(fileDiv);fileProgressBars[info.file]fileDiv.querySelector(.progress-fill);}// 更新进度条fileProgressBars[info.file].style.width${info.progress}%;constmb(info.loaded/1024/1024).toFixed(2);consttotalMb(info.total/1024/1024).toFixed(2);statusDiv.textContent${info.file}:${mb}/${totalMb}MB;}if(info.statusready){statusDiv.textContentModel ready!;}}});更多进度跟踪示例请参阅上文本节中的示例。自定义配置覆盖模型的默认配置import{pipeline}fromhuggingface/transformers;constpipeawaitpipeline(text-generation,model-id,{config:{max_length:512,temperature:0.8,// ... 其他配置选项}});使用场景覆盖默认的生成参数调整模型特定设置在不修改模型文件的情况下测试不同的配置模型加载选项缓存目录指定下载模型的缓存位置// Node.js自定义缓存位置constpipeawaitpipeline(sentiment-analysis,model-id,{cache_dir:./my-custom-cache});默认行为如果未指定使用env.cacheDir默认./.cache仅在env.useFSCache true时生效Node.js浏览器缓存使用 Cache API通过env.cacheKey配置仅本地文件禁止任何网络请求constpipeawaitpipeline(sentiment-analysis,model-id,{local_files_only:true});使用场景离线应用隔离air-gapped环境使用预下载模型进行测试打包模型的部署环境重要说明模型必须已经缓存或可在本地获取如果在本地找不到模型将抛出错误需要env.allowLocalModels true模型版本revision指定特定的模型版本git 分支、标签或提交constpipeawaitpipeline(sentiment-analysis,model-id,{revision:v1.0.0// 使用特定版本});// 或使用分支constpipeawaitpipeline(sentiment-analysis,model-id,{revision:experimental});// 或使用提交哈希constpipeawaitpipeline(sentiment-analysis,model-id,{revision:abc123def456});默认值main最新版本使用场景生产环境锁定到稳定版本测试实验性功能使用特定模型版本复现结果使用开发中的模型重要说明仅适用于远程模型Hugging Face Hub对本地文件路径无效每个版本单独缓存模型子文件夹指定模型仓库内的子文件夹constpipeawaitpipeline(sentiment-analysis,model-id,{subfolder:onnx// 默认onnx});默认值onnx使用场景自定义模型仓库结构同一仓库中的多个模型变体组织偏好模型文件名指定自定义模型文件名不含.onnx扩展名constpipeawaitpipeline(text-generation,model-id,{model_file_name:decoder_model_merged});// 加载decoder_model_merged.onnx使用场景文件名不标准的模型选择特定的模型变体文件分离的编码器-解码器模型注意目前仅对仅编码器或仅解码器模型有效。设备与性能选项设备选择选择运行模型的位置// 在 CPU 上运行WASM - 默认constpipeawaitpipeline(sentiment-analysis,model-id,{device:wasm});// 在 GPU 上运行WebGPUconstpipeawaitpipeline(sentiment-analysis,model-id,{device:webgpu});常见设备wasm- WebAssemblyCPU兼容性最好webgpu- WebGPU支持的运行时中的 GPU 加速cpu- CPUgpu- 自动检测 GPUcuda- NVIDIA CUDA带 GPU 的 Node.js完整列表请参阅 devices.js 源码。按组件选择设备对于具有多个组件的模型编码器-解码器、视觉-编码器-解码器等constpipeawaitpipeline(automatic-speech-recognition,model-id,{device:{encoder:webgpu,// 在 GPU 上运行编码器decoder:wasm// 在 CPU 上运行解码器}});WebGPU 要求支持 WebGPU 的运行时浏览器、Node.js、Bun 或 Deno兼容的硬件/驱动程序栈足够的 GPU 内存数据类型量化控制模型精度和大小// 全精度最大、最准确constpipeawaitpipeline(sentiment-analysis,model-id,{dtype:fp32});// 半精度均衡constpipeawaitpipeline(sentiment-analysis,model-id,{dtype:fp16});// 8 位量化更小、更快constpipeawaitpipeline(sentiment-analysis,model-id,{dtype:q8});// 4 位量化最小、最快constpipeawaitpipeline(sentiment-analysis,model-id,{dtype:q4});常见数据类型fp32- 32 位浮点全精度fp16- 16 位浮点半精度q8- 8 位量化良好平衡q4- 4 位量化最大压缩int8- 8 位整数uint8- 8 位无符号整数完整列表请参阅 dtypes.js 源码。按组件设置数据类型constpipeawaitpipeline(automatic-speech-recognition,model-id,{dtype:{encoder:fp32,// 编码器使用全精度decoder:q8// 解码器使用量化}});权衡数据类型模型大小速度精度使用场景fp32最大最慢最高研究、最高质量fp16中等中等高生产环境、GPU 推理q8小快良好生产环境、CPU 推理q4最小最快可接受边缘设备、实时应用外部数据格式对于 2GB 的模型ONNX 使用外部数据格式// 自动检测并加载外部数据constpipeawaitpipeline(text-generation,large-model-id,{use_external_data_format:true});// 指定外部数据块数量constpipeawaitpipeline(text-generation,large-model-id,{use_external_data_format:5// 加载 5 个数据块model.onnx_data_0 到 _4});工作原理 2GB 的模型会将权重拆分到多个文件中主文件model.onnx仅结构数据文件model.onnx_data或model.onnx_data_0、model.onnx_data_1等默认行为false- 无外部数据模型 2GBtrue- 自动加载外部数据number- 加载指定数量的外部数据块最大数据块数100由MAX_EXTERNAL_DATA_CHUNKS定义按组件设置外部数据constpipeawaitpipeline(text-generation,large-model-id,{use_external_data_format:{encoder:true,decoder:3// 解码器有 3 个外部数据块}});会话选项高级 ONNX Runtime 配置constpipeawaitpipeline(sentiment-analysis,model-id,{session_options:{executionProviders:[webgpu,wasm],graphOptimizationLevel:all,enableCpuMemArena:true,enableMemPattern:true,executionMode:sequential,logSeverityLevel:2,logVerbosityLevel:0}});常见会话选项选项描述默认值executionProviders有序的执行提供程序列表[wasm]graphOptimizationLevel图优化disabled、basic、extended、allallenableCpuMemArena启用 CPU 内存池以加快内存分配trueenableMemPattern启用内存模式优化trueexecutionModesequential或parallelsequentiallogSeverityLevel0详细1信息2警告3错误4致命2freeDimensionOverrides覆盖动态维度例如{ batch_size: 1 }-使用场景针对特定硬件微调性能调试模型执行问题覆盖动态形状控制内存使用常见配置模式开发环境带进度跟踪的快速迭代import{pipeline}fromhuggingface/transformers;constpipeawaitpipeline(sentiment-analysis,null,{progress_callback:(info){if(info.statusprogress_total){console.log(Total:${info.progress.toFixed(1)}%);}}});生产环境GPU使用 WebGPU 和 fp16 以获得更好的性能constpipeawaitpipeline(sentiment-analysis,model-id,{device:webgpu,dtype:fp16});生产环境CPU使用量化以获得更小的体积和更快的 CPU 推理constpipeawaitpipeline(sentiment-analysis,model-id,{dtype:q8// 或使用 q4 进一步减小体积});离线/本地禁止网络请求仅使用本地模型import{pipeline,env}fromhuggingface/transformers;env.allowLocalModelstrue;env.localModelPath./models/;constpipeawaitpipeline(sentiment-analysis,model-id,{local_files_only:true});按组件设置对于编码器-解码器模型分别配置每个组件constpipeawaitpipeline(automatic-speech-recognition,model-id,{device:{encoder:webgpu,decoder:wasm},dtype:{encoder:fp16,decoder:q8}});相关文档配置参考 - 使用env对象进行环境配置文本生成指南 - 文本生成选项与流式输出模型架构 - 支持的模型与选择建议主技能指南 - Transformers.js 入门最佳实践进度回调对大型模型使用progress_callback显示下载进度量化CPU 推理使用q8或q4以减小体积并提高速度设备选择可用时使用webgpu以获得更好的性能离线优先生产环境中使用local_files_only: true避免运行时下载版本锁定使用revision锁定模型版本以实现可复现的部署内存管理完成后始终使用pipe.dispose()释放管道本文档涵盖了pipeline()函数的全部可用选项。有关环境级配置远程主机、全局缓存设置、WASM 路径请参阅 配置参考。
网站建设高端定制企业官网