新闻详情

新闻详情

首页 / 资讯中心 / 详情

基于 Docker 的 Thanos 部署:TaoToken 统一 Key 打通 MinIO 与 Prometheus 长期存储

发布时间:2026/9/29 14:53:56来源:尧图网络
基于 Docker 的 Thanos 部署:TaoToken 统一 Key 打通 MinIO 与 Prometheus 长期存储
1. 为什么单机 Prometheus 撑不住长期指标Thanos 到底解决什么问题Prometheus 本地 TSDB 默认保留 15 天超过就删。你想查上个月的接口 P99或者对比两个集群同一时段的 QPS单机 Prometheus 直接告诉你「no data」。更麻烦的是多集群场景每个集群一套 PrometheusGrafana 要配 N 个数据源跨集群聚合查询基本靠人肉拼。Thanos 的思路很直接让 Prometheus 继续干它擅长的采集和短期存储Sidecar 把已经落盘的 2 小时 block 上传到对象存储MinIOStore Gateway 负责从对象存储里读历史数据Query 层把实时数据和历史数据合并成一个统一查询入口。Compactor 在后台做降采样和 block 合并把长期存储的成本压下来。这套架构适合谁我总结三类一是指标保留周期要求超过 30 天的团队二是多集群/多 Prometheus 实例需要全局查询视图的三是已经在用 Grafana 但被数据源切换折磨的。如果你只是单机跑个监控看 CPUThanos 属于杀鸡用牛刀。本文用 Docker Compose 把 Thanos 的 Sidecar、Store、Query、Compact 四个核心组件串起来对象存储用 MinIO采集端用 Prometheus最后在 Grafana 里验证查询结果。所有配置可直接复制IP 按你自己的环境替换。这里有个容易被忽略的点Thanos 各组件之间靠 gRPC 通信Sidecar 和 Store 都要暴露 gRPC 端口给 Query。如果你用 Docker 默认 bridge 网络容器间跨主机通信会踩坑所以下面统一用--network host省掉端口映射的麻烦。另外Thanos 对 Prometheus 的external_labels有硬性要求。多副本 Prometheus 必须有不同的replica标签否则 Query 去重会出问题。这个在配置 Prometheus 时就要改不能等部署完再补。2. TaoToken 统一 Key 在 Thanos 链路里的定位与前置准备Thanos 本身不涉及大模型调用但你在做监控告警智能化、Grafana 异常检测、或者用 AI 辅助分析指标趋势时就需要一个统一的模型接入层。TaoToken 在这里的角色是把不同模型供应商的 Key 收敛成一个避免在 Grafana 插件、告警机器人、运维脚本里到处散落 API Key。我试过在告警通知链路里接一个模型做告警摘要之前每个 webhook 脚本里硬编码 Key换模型就要改一堆文件。后来统一走 TaoToken 的 API 端点脚本里只留一个 Key模型切换在控制台改就行。前置准备分两步。第一步拿到 TaoToken 的 API Key。访问控制台创建# 控制台地址创建 API Key https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentconsole创建后复制 Key格式类似sk-xxxxxxxx。这个 Key 后面会用在告警摘要脚本和 Grafana 的 AI 面板里。第二步确认你的模型 ID。TaoToken 的 API 兼容 OpenAI 格式Base URL 是https://taotoken.net/api注意这个地址不加 UTM 参数直接用于代码里的base_url。模型 ID 在模型对话页面可以查到# 模型对话查看可用模型 ID https://taotoken.net/models?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentmodel_chat如果你只是做 Thanos 部署TaoToken 不是必须的。但如果你想让告警邮件里自动带上「这个指标异常可能是什么原因」的 AI 分析或者用 Coding Plan 让 AI 帮你写 Thanos 查询语句那就需要提前把 Key 准备好。# Coding Plan长期编码/Agent 场景 https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentcoding_plan接入文档在这里里面有各语言的调用示例# 接入文档 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_contentdoc3. 可复制的 docker-compose.yml 与 Thanos 各组件配置这一节是核心。我把 MinIO、Prometheus、Thanos Sidecar/Store/Query/Compact 全部写进一个 compose 文件你改一下 IP 和路径就能跑。先建目录结构mkdir -p /data/thanos/{minio,prometheus,thanos,grafana} mkdir -p /data/thanos/prometheus/{data,conf/rules} mkdir -p /data/thanos/thanos/{conf,store,compact} chown -R 65534:65534 /data/thanos/prometheus/dataMinIO 的 bucket 配置文件bucket_config.yamltype: S3 config: bucket: thanos endpoint: 192.168.11.193:9000 access_key: admin secret_key: admin123456 insecure: true signature_version2: false注意insecure: true是因为 MinIO 没上 TLS生产环境建议配证书。signature_version2保持 falseThanos 默认用 v4 签名。Prometheus 配置prometheus.yml重点是external_labelsglobal: scrape_interval: 30s evaluation_interval: 30s external_labels: region: GuangZhou replica: A rule_files: - /etc/prometheus/rules/*.rules scrape_configs: - job_name: prometheus static_configs: - targets: [localhost:9090] - job_name: thanos_sidecar static_configs: - targets: [192.168.11.193:19191]replica: A这个标签很关键。如果你有两套 Prometheus 采同样的目标Query 层靠--query.replica-labelreplica去重没有这个标签会返回重复数据。docker-compose.yml 完整内容version: 3.8 services: minio: image: minio/minio:latest container_name: minio restart: always network_mode: host environment: MINIO_ROOT_USER: admin MINIO_ROOT_PASSWORD: admin123456 MINIO_PROMETHEUS_AUTH_TYPE: public volumes: - /data/thanos/minio/data:/data - /etc/localtime:/etc/localtime:ro command: server /data --console-address :9001 prometheus: image: prom/prometheus:v2.28.0 container_name: prometheus restart: always network_mode: host volumes: - /data/thanos/prometheus/conf/prometheus.yml:/etc/prometheus/prometheus.yml - /data/thanos/prometheus/conf/rules:/etc/prometheus/rules - /data/thanos/prometheus/data:/data/prometheus/data - /etc/localtime:/etc/localtime:ro command: - --config.file/etc/prometheus/prometheus.yml - --storage.tsdb.path/data/prometheus/data - --storage.tsdb.retention30d - --storage.tsdb.min-block-duration2h - --storage.tsdb.max-block-duration2h - --web.enable-lifecycle - --web.enable-admin-api thanos-sidecar: image: quay.io/thanos/thanos:v0.28.0 container_name: thanos-sidecar restart: always network_mode: host volumes: - /data/thanos/prometheus/data:/data/prometheus/data - /data/thanos/thanos/conf/bucket_config.yaml:/bucket_config.yaml - /etc/localtime:/etc/localtime:ro command: - sidecar - --tsdb.path/data/prometheus/data - --prometheus.urlhttp://192.168.11.193:9090 - --objstore.config-file/bucket_config.yaml - --http-address0.0.0.0:19191 - --grpc-address0.0.0.0:19090 thanos-store: image: quay.io/thanos/thanos:v0.28.0 container_name: thanos-store restart: always network_mode: host volumes: - /data/thanos/thanos/store:/var/thanos/store - /data/thanos/thanos/conf/bucket_config.yaml:/bucket_config.yaml - /etc/localtime:/etc/localtime:ro command: - store - --data-dir/var/thanos/store - --objstore.config-file/bucket_config.yaml - --http-address0.0.0.0:29191 - --grpc-address0.0.0.0:29090 - --index-cache-size1GB - --chunk-pool-size8GB thanos-query: image: quay.io/thanos/thanos:v0.28.0 container_name: thanos-query restart: always network_mode: host volumes: - /data/thanos/thanos/conf/store.yaml:/store.yaml - /etc/localtime:/etc/localtime:ro command: - query - --http-address0.0.0.0:19192 - --grpc-address0.0.0.0:19091 - --store.sd-files/store.yaml - --query.replica-labelreplica thanos-compact: image: quay.io/thanos/thanos:v0.28.0 container_name: thanos-compact restart: always network_mode: host volumes: - /data/thanos/thanos/compact:/var/thanos/compact - /data/thanos/thanos/conf/bucket_config.yaml:/bucket_config.yaml - /etc/localtime:/etc/localtime:ro command: - compact - --data-dir/var/thanos/compact - --objstore.config-file/bucket_config.yaml - --http-address0.0.0.0:19193 - --waitStore 的服务发现文件store.yaml- targets: - 192.168.11.193:19090 - 192.168.11.193:29090这里19090是 Sidecar 的 gRPC 端口29090是 Store 的 gRPC 端口。Query 通过这个文件发现所有数据源。启动顺序有讲究先起 MinIO创建 bucket再起 Prometheus最后起 Thanos 组件。因为 Sidecar 启动时会检查 bucket 是否存在bucket 没建好会报错退出。cd /data/thanos docker-compose up -d minio # 等 MinIO 起来后浏览器访问 http://192.168.11.193:9001 创建 bucket thanos docker-compose up -d prometheus docker-compose up -d thanos-sidecar thanos-store thanos-query thanos-compact4. 验证请求从 MinIO bucket 到 Grafana 查询结果核对部署完不能只看容器起没起要验证数据真的从 Prometheus 流到了 MinIO再从 Query 查得出来。第一步确认 Sidecar 上传成功。访问 Sidecar 的 HTTP 接口curl -s http://192.168.11.193:19191/api/v1/status | jq返回里关注objstore部分如果看到status: ok说明 bucket 连接正常。再看 MinIO 控制台thanosbucket 下应该出现01H...开头的 block 目录里面有chunks和index文件。第二步验证 Store Gateway 能读到历史 blockcurl -s http://192.168.11.193:29191/api/v1/stores | jq返回的stores数组里应该有thanos-store自己的信息minTime和maxTime覆盖你上传的 block 时间范围。第三步Query 层全局查询。访问http://192.168.11.193:19192/graph在查询框输入up{jobprometheus}点 Execute应该返回1。再查一个历史时间点的数据比如 2 小时前up{jobprometheus}[2h]如果 Query 同时连了 Sidecar 和 Store这个查询会合并实时数据和历史数据。你可以在 Query 的 UI 上看到Stores列表确认两个 gRPC 端点都在。第四步Grafana 数据源配置。Grafana 启动后添加 Prometheus 类型数据源URL 填http://192.168.11.193:19192也就是 Query 的地址不是 Prometheus 的 9090。这一步很多人配错配成 Prometheus 地址就只能查实时数据历史 block 查不到。docker run -d \ --name grafana \ --restartalways \ --network host \ -e GF_SECURITY_ADMIN_PASSWORDGrafana2021 \ -v /data/thanos/grafana:/var/lib/grafana \ grafana/grafana:9.5.0Grafana 里导入 Thanos 官方 dashboardID 是14605。导入后在变量里选datasource为刚配的 Query 数据源面板应该能显示 Sidecar 上传的 block 数量和 Store 的查询延迟。第五步核对查询结果。在 Grafana Explore 里分别查prometheus_tsdb_head_samples_appended_total和thanos_store_bucket_blocks_count前者来自 Prometheus 实时数据后者来自 Store 读 MinIO 的统计。两个都有数据说明整条链路通了。如果你在告警链路里接了 TaoToken 做 AI 摘要可以在 Grafana 的 webhook 里加一段脚本import requests def summarize_alert(alert_text): resp requests.post( https://taotoken.net/api/v1/chat/completions, headers{Authorization: Bearer sk-你的Key}, json{ model: 你的模型ID, messages: [{role: user, content: f用一句话总结这个告警{alert_text}}] } ) return resp.json()[choices][0][message][content]注意base_url是https://taotoken.net/api不要加 UTM 参数。模型 ID 从模型对话页面复制。5. 本篇常见报错排查401、local proxy failed、reading choices、OAuth部署 Thanos 和接 TaoToken 时我踩过的坑集中在几个报错上逐个说。报错一Sidecar 启动后日志报no such bucket或Access Denied这是 MinIO bucket 没建或者 Key 不对。检查bucket_config.yaml里的access_key和secret_key是否和 MinIO 环境变量一致。另外endpoint不要带http://前缀只写IP:端口。如果 MinIO 是单节点insecure: true必须加。报错二Query 查询返回local proxy failed这个报错通常出现在 Query 连不上 Store 的 gRPC 端口。检查store.yaml里的地址是否可达docker exec thanos-query wget -qO- http://192.168.11.193:29090如果连不上确认 Store 容器是否在跑以及--network host是否生效。用 bridge 网络时容器间要用容器名而不是 IP。报错三TaoToken 调用返回 401{error: {message: Invalid API key, type: invalid_request_error}}检查 Key 是否复制完整有没有多余空格。另外确认请求头是Authorization: Bearer sk-xxx不是x-api-key。TaoToken 兼容 OpenAI 格式用 Bearer 认证。报错四reading choices相关错误{error: reading choices: unexpected end of JSON input}这个一般是响应体为空或者不是 JSON。先确认base_url写的是https://taotoken.net/api不是https://taotoken.net。路径要带/v1/chat/completions。如果用的是某些 SDK它会自动拼/v1那 base_url 就写到/api为止。报错五OAuth 相关报错如果你用 Claude Code 或 Codex 接入 TaoToken遇到 OAuth 报错检查配置文件。Claude Code 的配置在~/.claude/settings.json{ env: { ANTHROPIC_BASE_URL: https://taotoken.net/api, ANTHROPIC_API_KEY: sk-你的Key } }Codex 的auth.json在~/.codex/auth.json{ OPENAI_API_KEY: sk-你的Key, OPENAI_BASE_URL: https://taotoken.net/api }三件套必须齐全Base URL、Key、Model ID。缺一个就会报 OAuth 或认证失败。Model ID 从模型对话页面查不要自己猜。报错六Compactor 报bucket not found或一直waitCompactor 启动时如果 bucket 里没有 block会一直等。这是正常的等 Sidecar 上传第一个 block 后就会开始工作。如果超过 2 小时还没动静检查 Sidecar 的--min-time参数默认是 0不用改。6. 从部署到长期运行Thanos 与 TaoToken 的配合建议Thanos 跑起来只是开始长期运行有几个参数要调。Compactor 的--retention.resolution-raw默认 0表示永久保留原始数据。如果你只想留 90 天加上--retention.resolution-raw90d。降采样后的 5 分钟精度数据可以留更久用--retention.resolution-5m365d。Store Gateway 的内存占用和 block 数量成正比。--index-cache-size1GB和--chunk-pool-size8GB是我在 500 万 series 规模下的配置你可以根据实际数据量调整。如果 Store 频繁 OOM先加--index-cache-size再考虑加内存。Query 层的--query.replica-labelreplica必须和 Prometheus 的external_labels对应。如果你有多个副本Query 会自动去重。没有副本就留空但标签名要一致。TaoToken 在监控链路里的价值我总结两个实际场景。一是告警降噪Alertmanager 发出来的告警批量丢给模型做聚合把「同一台机器 10 个指标异常」合并成一条「机器 X 疑似磁盘 IO 瓶颈」。二是 Grafana 面板的 AI 解读在 dashboard 里加一个 Text 面板用 JS 调 TaoToken API把当前时间段的指标趋势翻译成自然语言。这两个场景都不需要把 TaoToken 写进 Thanos 的核心链路而是作为旁路增强。核心链路保持稳定AI 层挂了不影响监控。最后说一个运维细节Thanos 的 block 上传是 Sidecar 做的但 Sidecar 只上传已经落盘的 2 小时 block。如果你重启 Prometheus正在写入的 block 不会上传等它落盘后 Sidecar 会自动补传。所以不要手动去删 Prometheus 的 data 目录否则会丢数据。验证整条链路是否健康我习惯用这个命令curl -s http://192.168.11.193:19192/api/v1/stores | jq .stores[] | {name, lastCheck, minTime, maxTime}返回里每个 store 的lastCheck应该是几秒前minTime和maxTime覆盖你的数据范围。如果某个 store 的lastCheck是几分钟前说明 gRPC 连接断了去查那个组件的日志。这套配置我在测试环境跑了三个月MinIO 存了 200GB 左右的 blockQuery 查询 30 天范围的数据响应在 2 秒内。Compactor 每天凌晨做一次降采样CPU 峰值 40%。如果你数据量更大把 Compactor 单独放一台机器别和 Store 抢资源。
网站建设高端定制企业官网
RELATED

相关资讯

更多精彩内容,欢迎继续阅读

较早相关资讯

最新相关资讯

PX4无人机避障实战:3DVFH*算法与伴侣计算机架构解析 2026/9/29 15:46:16

PX4无人机避障实战:3DVFH*算法与伴侣计算机架构解析

1. 整体设计与方案选型思路1.1 为什么是“伴侣计算机飞控”的分工模式做PX4避障,很多人第一个反应是在飞控固件里加逻辑,但真正跑了几个版本你会发现,这条路走起来非常别扭。飞控的实时性很强,姿态环、位置环都在1kHz甚至更快的节…

阅读更多 →
LlamaFactory+LoRA微调Qwen2.5:法律与医疗领域模型实战指南 2026/9/29 15:46:16

LlamaFactory+LoRA微调Qwen2.5:法律与医疗领域模型实战指南

简介:一份面向AI研发、自然语言处理及法律/医疗智能化应用开发者的技术文档,讲解如何借助LlamaFactory对Qwen2.5大模型实施高效微调。文档基于LoRA、QLoRA等参数高效微调技术,结合多模型兼容、可视化界面与全流程监控特性,系统介绍…

阅读更多 →
从投诉数据看钓鱼攻击新趋势:防御机制优化实战 2026/9/29 15:46:16

从投诉数据看钓鱼攻击新趋势:防御机制优化实战

刚把2026年2月的钓鱼网站投诉数据跑完,趁着热乎劲儿把这一个月的复盘结论整理出来。做安全运营的都知道,投诉数据是最真实的“用户侧情报”——攻击者是不是换了套路,哪些仿冒对象又被盯上了,防御策略哪里漏了风,翻一遍…

阅读更多 →
YOLOv11提速5倍:剪枝与INT8量化完整实践指南 2026/9/29 15:46:16

YOLOv11提速5倍:剪枝与INT8量化完整实践指南

简介:YOLOv11作为单阶段目标检测代表,兼顾速度与精度,但复杂模型部署在边缘设备时仍面临高算力、高内存压力。为应对这一挑战,这份实战PDF围绕剪枝、量化两条主线,给出从原理到落地的一条龙方案,适合算法工…

阅读更多 →
openFuyao v25.09生产环境部署:GPU/ARM异构算力调度实践与踩坑记录 2026/9/29 15:46:16

openFuyao v25.09生产环境部署:GPU/ARM异构算力调度实践与踩坑记录

先把话说明白:如果你的团队只有三五台机器,手动分配算力完全够用,不用折腾任何调度平台。但当你手里同时有 x86 服务器、ARM 边缘盒子、NVIDIA 显卡、NPU 推理棒,资源又分散在不同项目组的时候,手动分配就开始失控了—…

阅读更多 →
PyTorch神经网络入门:自动求导、动态图与CNN实战 2026/9/29 15:46:09

PyTorch神经网络入门:自动求导、动态图与CNN实战

1. 先用一句话说清楚:PyTorch 里的神经网络到底在做什么很多人第一次接触 PyTorch,是因为听说它"好用、上手快",但真正打开官方文档,看到nn.Module、autograd、backward()这一堆名词,瞬间就懵了。我用一句话…

阅读更多 →

今日资讯

本周资讯

本月资讯

看完文章仍有疑问?

联系尧图顾问,获取一对一建站咨询

立即免费咨询 📞 400-888-8888
📞 ✉