新闻详情

新闻详情

首页 / 资讯中心 / 详情

第3讲:LLM 调用监控

发布时间:2026/10/1 21:55:39来源:尧图网络
第3讲:LLM 调用监控
一、为什么需要专门的 LLM 监控LLM 调用是 AI 应用中最昂贵、最不可控的环节。一个典型的客服系统每天可能调用数十万次 LLM每次调用都涉及延迟几百毫秒到几十秒不等成本按 Token 计费积少成多质量输出不稳定可能产生幻觉异常超时、限流、截断、空回复没有监控你无法回答这些问题今天花了多少钱哪个用户花得最多哪个模型的延迟最高哪个时段最慢有多少请求被截断了有多少产生了空回复二、监控架构┌─────────────────────────────────────────────────────────────┐ │ LLM 调用监控架构 │ │ │ │ 应用层 │ │ ┌──────────────────────────────────────────────────┐ │ │ │ LLM Client 中间件 (自动埋点) │ │ │ │ ├─ BeforeCall: 记录请求时间、Prompt Token 预估 │ │ │ │ ├─ AfterCall: 记录延迟、Token 消耗、错误信息 │ │ │ │ └─ OnError: 记录错误类型、重试次数 │ │ │ └──────────────────────┬───────────────────────────┘ │ │ │ │ │ 采集层 ▼ │ │ ┌──────────────────────────────────────────────────┐ │ │ │ LLMMetricCollector │ │ │ │ ├─ Per-Call 指标 (延迟、Token、成本) │ │ │ │ ├─ Per-Model 聚合 (平均延迟、P99、错误率) │ │ │ │ └─ Per-User 聚合 (调用次数、总成本) │ │ │ └──────────────────────┬───────────────────────────┘ │ │ │ │ │ 分析层 ▼ │ │ ┌──────────────────────────────────────────────────┐ │ │ │ 实时分析 告警引擎 │ │ │ │ ├─ 延迟突增告警 (P99 5s) │ │ │ │ ├─ 成本超标告警 (单用户 $100/天) │ │ │ │ ├─ 异常检测 (截断率 10%、空回复率 5%) │ │ │ │ └─ 模型退化告警 (同模型延迟上涨 50%) │ │ │ └──────────────────────┬───────────────────────────┘ │ │ │ │ │ 展示层 ▼ │ │ ┌──────────────────────────────────────────────────┐ │ │ │ Grafana 仪表盘 告警通知 │ │ │ │ ├─ 实时 QPS / 延迟 / 成本 │ │ │ │ ├─ 模型对比面板 │ │ │ │ ├─ 用户 TopN │ │ │ │ └─ 异常事件列表 │ │ │ └──────────────────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────┘三、完整代码实现package main import ( context encoding/json fmt math sort strings sync time ) // // 1. 核心数据结构 // // LLMCallRecord 单次 LLM 调用的完整记录 type LLMCallRecord struct { CallID string json:call_id Model string json:model Provider string json:provider UserID string json:user_id SessionID string json:session_id StartTime time.Time json:start_time EndTime time.Time json:end_time DurationMs int64 json:duration_ms TTFTMs int64 json:ttft_ms // Time to first token PromptTokens int64 json:prompt_tokens CompletionTokens int64 json:completion_tokens TotalTokens int64 json:total_tokens CostUSD float64 json:cost_usd Temperature float64 json:temperature MaxTokens int64 json:max_tokens Truncated bool json:truncated EmptyResponse bool json:empty_response Error string json:error,omitempty RetryCount int json:retry_count Status string json:status // success / error / timeout / rate_limited Tags map[string]string json:tags } // LLMMetrics 聚合指标 type LLMMetrics struct { TotalCalls int64 json:total_calls SuccessCalls int64 json:success_calls ErrorCalls int64 json:error_calls TimeoutCalls int64 json:timeout_calls RateLimitedCalls int64 json:rate_limited_calls TruncatedCalls int64 json:truncated_calls EmptyResponses int64 json:empty_responses AvgDurationMs float64 json:avg_duration_ms P50DurationMs float64 json:p50_duration_ms P95DurationMs float64 json:p95_duration_ms P99DurationMs float64 json:p99_duration_ms AvgTTFTMs float64 json:avg_ttft_ms TotalPromptTokens int64 json:total_prompt_tokens TotalCompletionTokens int64 json:total_completion_tokens TotalTokens int64 json:total_tokens TotalCostUSD float64 json:total_cost_usd AvgCostPerCall float64 json:avg_cost_per_call ErrorRate float64 json:error_rate TruncatedRate float64 json:truncated_rate EmptyRate float64 json:empty_rate } // // 2. 模型定价表 // type ModelPricing struct { InputPricePer1M float64 // $ per 1M input tokens OutputPricePer1M float64 // $ per 1M output tokens } var defaultPricing map[string]ModelPricing{ gpt-4o: {InputPricePer1M: 2.50, OutputPricePer1M: 10.00}, gpt-4o-mini: {InputPricePer1M: 0.15, OutputPricePer1M: 0.60}, gpt-4-turbo: {InputPricePer1M: 10.00, OutputPricePer1M: 30.00}, claude-3-opus: {InputPricePer1M: 15.00, OutputPricePer1M: 75.00}, claude-3-sonnet: {InputPricePer1M: 3.00, OutputPricePer1M: 15.00}, claude-3-haiku: {InputPricePer1M: 0.25, OutputPricePer1M: 1.25}, deepseek-chat: {InputPricePer1M: 0.14, OutputPricePer1M: 0.55}, gemini-pro: {InputPricePer1M: 0.125, OutputPricePer1M: 0.375}, } func calculateCost(model string, promptTokens, completionTokens int64) float64 { pricing, ok : defaultPricing[model] if !ok { // 未知模型按最低价估算 pricing ModelPricing{InputPricePer1M: 0.15, OutputPricePer1M: 0.60} } promptCost : float64(promptTokens) * pricing.InputPricePer1M / 1_000_000 completionCost : float64(completionTokens) * pricing.OutputPricePer1M / 1_000_000 return promptCost completionCost } func extractProvider(model string) string { if strings.Contains(model, gpt) || strings.Contains(model, o1) { return openai } if strings.Contains(model, claude) { return anthropic } if strings.Contains(model, gemini) { return google } if strings.Contains(model, deepseek) { return deepseek } return unknown } // // 3. LLM Client 中间件自动埋点 // // LLMClient LLM 客户端接口 type LLMClient interface { Chat(ctx context.Context, req *ChatRequest) (*ChatResponse, error) } type ChatRequest struct { Model string Messages []Message Temperature float64 MaxTokens int64 UserID string SessionID string } type Message struct { Role string json:role Content string json:content } type ChatResponse struct { Content string PromptTokens int64 CompletionTokens int64 TTFTMs int64 Truncated bool } // MonitoredLLMClient 带监控的 LLM 客户端包装器 type MonitoredLLMClient struct { inner LLMClient collector *LLMMetricCollector config MonitorConfig } type MonitorConfig struct { RecordPayload bool // 是否记录请求/响应内容注意隐私 SlowThresholdMs int64 // 慢调用阈值超过记录警告 CostAlertUSD float64 // 单次调用成本告警阈值 MaxRetries int // 最大重试次数 TimeoutMs int64 // 超时时间 } func NewMonitoredLLMClient(inner LLMClient, collector *LLMMetricCollector, config MonitorConfig) *MonitoredLLMClient { if config.MaxRetries 0 { config.MaxRetries 3 } if config.TimeoutMs 0 { config.TimeoutMs 30000 // 30秒 } if config.SlowThresholdMs 0 { config.SlowThresholdMs 5000 // 5秒 } return MonitoredLLMClient{ inner: inner, collector: collector, config: config, } } // Chat 带监控的调用 func (c *MonitoredLLMClient) Chat(ctx context.Context, req *ChatRequest) (*ChatResponse, error) { callID : fmt.Sprintf(call-%x, time.Now().UnixNano()) startTime : time.Now() // 预估 Prompt Token粗略估计1 token ≈ 4 字符 estimatedPromptTokens : estimateTokens(req.Messages) var lastErr error var resp *ChatResponse // 重试逻辑 for attempt : 0; attempt c.config.MaxRetries; attempt { if attempt 0 { // 退避等待 backoff : time.Duration(math.Pow(2, float64(attempt))) * 100 * time.Millisecond time.Sleep(backoff) } resp, lastErr c.inner.Chat(ctx, req) if lastErr nil { break } // 判断是否值得重试 if isRetryable(lastErr) { continue } break } endTime : time.Now() durationMs : endTime.Sub(startTime).Milliseconds() // 构造调用记录 record : LLMCallRecord{ CallID: callID, Model: req.Model, Provider: extractProvider(req.Model), UserID: req.UserID, SessionID: req.SessionID, StartTime: startTime, EndTime: endTime, DurationMs: durationMs, Temperature: req.Temperature, MaxTokens: req.MaxTokens, RetryCount: attemptCount(lastErr, req.MaxRetries), Tags: map[string]string{ env: production, }, } if lastErr ! nil { record.Error lastErr.Error() record.Status classifyError(lastErr) } else { record.Status success record.PromptTokens estimatedPromptTokens record.CompletionTokens resp.CompletionTokens record.TotalTokens estimatedPromptTokens resp.CompletionTokens record.TTFTMs resp.TTFTMs record.Truncated resp.Truncated record.EmptyResponse len(resp.Content) 0 record.CostUSD calculateCost(req.Model, record.PromptTokens, record.CompletionTokens) } // 记录到采集器 c.collector.Record(record) // 告警检查 c.checkAlerts(record) // 慢调用日志 if durationMs c.config.SlowThresholdMs { fmt.Printf([SLOW] %s | %s | %dms | user%s\n, req.Model, callID[:12], durationMs, req.UserID) } return resp, lastErr } // // 4. 指标采集器 // type LLMMetricCollector struct { mu sync.RWMutex records []*LLMCallRecord modelAgg map[string]*LLMMetrics // 按模型聚合 userAgg map[string]*LLMMetrics // 按用户聚合 hourlyAgg map[int64]*LLMMetrics // 按小时聚合 maxRecords int // 最大保留记录数 } func NewLLMMetricCollector(maxRecords int) *LLMMetricCollector { if maxRecords 0 { maxRecords 100000 } return LLMMetricCollector{ records: make([]*LLMCallRecord, 0, maxRecords), modelAgg: make(map[string]*LLMMetrics), userAgg: make(map[string]*LLMMetrics), hourlyAgg: make(map[int64]*LLMMetrics), maxRecords: maxRecords, } } func (c *LLMMetricCollector) Record(record *LLMCallRecord) { c.mu.Lock() defer c.mu.Unlock() // 追加记录 c.records append(c.records, record) if len(c.records) c.maxRecords { c.records c.records[len(c.records)-c.maxRecords:] } // 更新模型聚合 c.updateAggregation(c.modelAgg, record.Model, record) // 更新用户聚合 c.updateAggregation(c.userAgg, record.UserID, record) // 更新小时聚合 hourKey : record.StartTime.Unix() / 3600 c.updateAggregation(c.hourlyAgg, fmt.Sprintf(%d, hourKey), record) } func (c *LLMMetricCollector) updateAggregation(aggMap map[string]*LLMMetrics, key string, record *LLMCallRecord) { agg, exists : aggMap[key] if !exists { agg LLMMetrics{} aggMap[key] agg } agg.TotalCalls switch record.Status { case success: agg.SuccessCalls case error: agg.ErrorCalls case timeout: agg.TimeoutCalls case rate_limited: agg.RateLimitedCalls } if record.Truncated { agg.TruncatedCalls } if record.EmptyResponse { agg.EmptyResponses } agg.TotalPromptTokens record.PromptTokens agg.TotalCompletionTokens record.CompletionTokens agg.TotalTokens record.TotalTokens agg.TotalCostUSD record.CostUSD agg.AvgDurationMs (agg.AvgDurationMs*float64(agg.TotalCalls-1) float64(record.DurationMs)) / float64(agg.TotalCalls) agg.AvgTTFTMs (agg.AvgTTFTMs*float64(agg.TotalCalls-1) float64(record.TTFTMs)) / float64(agg.TotalCalls) agg.AvgCostPerCall agg.TotalCostUSD / float64(agg.TotalCalls) agg.ErrorRate float64(agg.ErrorCallsagg.TimeoutCalls) / float64(agg.TotalCalls) * 100 agg.TruncatedRate float64(agg.TruncatedCalls) / float64(agg.TotalCalls) * 100 agg.EmptyRate float64(agg.EmptyResponses) / float64(agg.TotalCalls) * 100 } // GetModelMetrics 获取指定模型的聚合指标 func (c *LLMMetricCollector) GetModelMetrics(model string) *LLMMetrics { c.mu.RLock() defer c.mu.RUnlock() return c.modelAgg[model] } // GetUserMetrics 获取指定用户的聚合指标 func (c *LLMMetricCollector) GetUserMetrics(userID string) *LLMMetrics { c.mu.RLock() defer c.mu.RUnlock() return c.userAgg[userID] } // GetAllModelMetrics 获取所有模型的指标 func (c *LLMMetricCollector) GetAllModelMetrics() map[string]*LLMMetrics { c.mu.RLock() defer c.mu.RUnlock() result : make(map[string]*LLMMetrics) for k, v : range c.modelAgg { result[k] v } return result } // GetTopUsersByCost 获取消费最高的用户 func (c *LLMMetricCollector) GetTopUsersByCost(n int) []struct { UserID string Cost float64 } { c.mu.RLock() defer c.mu.RUnlock() type userCost struct { UserID string Cost float64 } var users []userCost for uid, metrics : range c.userAgg { users append(users, userCost{UserID: uid, Cost: metrics.TotalCostUSD}) } sort.Slice(users, func(i, j int) bool { return users[i].Cost users[j].Cost }) if n len(users) { n len(users) } result : make([]struct { UserID string Cost float64 }, n) for i : 0; i n; i { result[i].UserID users[i].UserID result[i].Cost users[i].Cost } return result } // // 5. 告警引擎 // type AlertRule struct { Name string Description string Severity string // critical / warning / info Check func(record *LLMCallRecord, metrics *LLMMetrics) bool } type AlertEngine struct { rules []AlertRule alertChan chan AlertEvent } type AlertEvent struct { RuleName string Severity string Message string Record *LLMCallRecord Timestamp time.Time } func NewAlertEngine(bufferSize int) *AlertEngine { return AlertEngine{ rules: make([]AlertRule, 0), alertChan: make(chan AlertEvent, bufferSize), } } func (e *AlertEngine) AddRule(rule AlertRule) { e.rules append(e.rules, rule) } func (e *AlertEngine) Evaluate(record *LLMCallRecord, metrics *LLMMetrics) { for _, rule : range e.rules { if rule.Check(record, metrics) { e.alertChan - AlertEvent{ RuleName: rule.Name, Severity: rule.Severity, Message: rule.Description, Record: record, Timestamp: time.Now(), } } } } func (e *AlertEngine) StartAlertHandler() { go func() { for alert : range e.alertChan { fmt.Printf(\n [%s] %s\n, alert.Severity, alert.RuleName) fmt.Printf( %s\n, alert.Message) fmt.Printf( 模型: %s | 用户: %s | 耗时: %dms\n, alert.Record.Model, alert.Record.UserID, alert.Record.DurationMs) if alert.Record.Error ! { fmt.Printf( 错误: %s\n, alert.Record.Error) } } }() } // // 6. 模拟 LLM Provider // type MockLLMProvider struct { name string latency time.Duration failRate float64 } func NewMockLLMProvider(name string, latency time.Duration, failRate float64) *MockLLMProvider { return MockLLMProvider{ name: name, latency: latency, failRate: failRate, } } func (p *MockLLMProvider) Chat(ctx context.Context, req *ChatRequest) (*ChatResponse, error) { // 模拟延迟 select { case -time.After(p.latency): case -ctx.Done(): return nil, fmt.Errorf(context canceled) } // 模拟随机失败 if randFloat() p.failRate { return nil, fmt.Errorf(rate_limit_exceeded: too many requests) } // 模拟截断 truncated : randFloat() 0.05 // 模拟空回复 empty : randFloat() 0.01 content : 这是一个模拟回复 if empty { content } return ChatResponse{ Content: content, PromptTokens: int64(len(req.Messages[0].Content) / 4), CompletionTokens: int64(len(content) / 4), TTFTMs: int64(p.latency.Milliseconds() / 2), Truncated: truncated, }, nil } // // 7. 辅助函数 // func estimateTokens(messages []Message) int64 { var totalChars int64 for _, msg : range messages { totalChars int64(len(msg.Content)) } // 粗略估计1 token ≈ 4 字符 return totalChars / 4 } func isRetryable(err error) bool { errStr : err.Error() retryable : []string{rate_limit, timeout, internal_error, server_error, 503} for _, keyword : range retryable { if strings.Contains(errStr, keyword) { return true } } return false } func classifyError(err error) string { errStr : err.Error() if strings.Contains(errStr, timeout) || strings.Contains(errStr, deadline) { return timeout } if strings.Contains(errStr, rate_limit) { return rate_limited } return error } func attemptCount(err error, maxRetries int) int { if err nil { return 0 } return maxRetries // 简化处理 } func randFloat() float64 { return float64(time.Now().UnixNano()%1000000) / 1000000.0 } // // 8. 主程序演示 // func main() { fmt.Println( 第3讲LLM 调用监控 \n) // 初始化采集器 collector : NewLLMMetricCollector(10000) // 初始化告警引擎 alertEngine : NewAlertEngine(100) alertEngine.AddRule(AlertRule{ Name: 高延迟告警, Description: 单次 LLM 调用延迟超过 10 秒, Severity: warning, Check: func(record *LLMCallRecord, metrics *LLMMetrics) bool { return record.DurationMs 10000 }, }) alertEngine.AddRule(AlertRule{ Name: 高成本告警, Description: 单次 LLM 调用成本超过 $0.05, Severity: info, Check: func(record *LLMCallRecord, metrics *LLMMetrics) bool { return record.CostUSD 0.05 }, }) alertEngine.AddRule(AlertRule{ Name: 频繁截断告警, Description: 模型截断率超过 10%, Severity: critical, Check: func(record *LLMCallRecord, metrics *LLMMetrics) bool { if metrics ! nil metrics.TotalCalls 100 { return metrics.TruncatedRate 10.0 } return false }, }) alertEngine.StartAlertHandler() // 创建多个模型客户端 models : []struct { name string latency time.Duration failRate float64 }{ {gpt-4o-mini, 800 * time.Millisecond, 0.02}, {gpt-4o, 2500 * time.Millisecond, 0.05}, {claude-3-haiku, 600 * time.Millisecond, 0.01}, {deepseek-chat, 900 * time.Millisecond, 0.03}, } clients : make([]*MonitoredLLMClient, 0) for _, m : range models { provider : NewMockLLMProvider(m.name, m.latency, m.failRate) client : NewMonitoredLLMClient(provider, collector, MonitorConfig{ SlowThresholdMs: 5000, CostAlertUSD: 0.05, MaxRetries: 2, TimeoutMs: 30000, }) clients append(clients, client) } // 模拟用户请求 users : []string{user_001, user_002, user_003, user_004} intents : []string{查询订单, 申请退款, 投诉建议, 产品咨询} fmt.Println(--- 模拟 50 次 LLM 调用 ---) for i : 0; i 50; i { client : clients[i%len(clients)] user : users[i%len(users)] intent : intents[i%len(intents)] req : ChatRequest{ Model: models[i%len(models)].name, Messages: []Message{{Role: user, Content: intent}}, Temperature: 0.7, MaxTokens: 1024, UserID: user, SessionID: fmt.Sprintf(session_%d, i/10), } resp, err : client.Chat(context.Background(), req) if err ! nil { fmt.Printf( [%d] %s | %s | ❌ %s\n, i1, user, req.Model, err.Error()) } else { status : ✅ if resp.Truncated { status ⚠️ } if resp.EmptyResponse { status ⬜ } fmt.Printf( [%d] %s | %s | %s | tokens%d\n, i1, user, req.Model, status, resp.PromptTokensresp.CompletionTokens) } time.Sleep(50 * time.Millisecond) } // 输出聚合报告 fmt.Println(\n--- 模型聚合指标 ---) for model, metrics : range collector.GetAllModelMetrics() { fmt.Printf(\n %s:\n, model) fmt.Printf( 调用次数: %d\n, metrics.TotalCalls) fmt.Printf( 成功率: %.1f%%\n, float64(metrics.SuccessCalls)/float64(metrics.TotalCalls)*100) fmt.Printf( 平均延迟: %.0fms\n, metrics.AvgDurationMs) fmt.Printf( 总 Token: %d (输入: %d / 输出: %d)\n, metrics.TotalTokens, metrics.TotalPromptTokens, metrics.TotalCompletionTokens) fmt.Printf( 总成本: $%.4f\n, metrics.TotalCostUSD) fmt.Printf( 平均成本/次: $%.6f\n, metrics.AvgCostPerCall) fmt.Printf( 截断率: %.1f%%\n, metrics.TruncatedRate) fmt.Printf( 空回复率: %.1f%%\n, metrics.EmptyRate) } // 用户消费排行 fmt.Println(\n--- 用户消费 Top 3 ---) topUsers : collector.GetTopUsersByCost(3) for i, u : range topUsers { fmt.Printf( %d. %s: $%.4f\n, i1, u.UserID, u.Cost) } // 汇总 fmt.Println(\n--- 总体统计 ---) totalCalls : int64(0) totalCost : 0.0 for _, metrics : range collector.GetAllModelMetrics() { totalCalls metrics.TotalCalls totalCost metrics.TotalCostUSD } fmt.Printf( 总调用次数: %d\n, totalCalls) fmt.Printf( 总成本: $%.4f\n, totalCost) fmt.Printf( 平均成本/次: $%.6f\n, totalCost/float64(totalCalls)) } // 确保 unused 变量不被编译报错 var _ sort.Slice四、关键指标解读4.1 延迟指标P50: 800ms ─── 一半请求比这个快 P95: 2500ms ─── 95% 的请求在这个时间内完成 P99: 5000ms ─── 1% 的请求非常慢需要关注P99 飙升的可能原因模型侧负载高共享实例网络抖动输入 Prompt 过长并发过高导致排队4.2 成本指标gpt-4o-mini: $0.0003/次 ← 性价比之王 gpt-4o: $0.0050/次 ← 贵 16 倍 claude-3-opus: $0.0150/次 ← 贵 50 倍优化思路简单问题走小模型复杂问题走大模型缓存重复请求缩短 Prompt减少输入 Token限制 Max Tokens减少输出 Token4.3 质量指标指标正常范围需关注告警截断率 5%5%-10% 10%空回复率 1%1%-3% 5%错误率 2%2%-5% 10%五、Client 中间件模式原始调用 带监控的调用 ┌─────────┐ ┌──────────────────┐ │ 业务代码 │ │ 业务代码 │ └────┬────┘ └───────┬──────────┘ │ 直接调用 │ 通过中间件 ▼ ▼ ┌─────────┐ ┌──────────────────┐ │ LLM API │ │ MonitoredLLMClient│ └─────────┘ │ ├─ 记录开始时间 │ │ ├─ 重试逻辑 │ │ ├─ 记录结束时间 │ │ ├─ 计算成本 │ │ ├─ 告警检查 │ │ └─ 调用真正的 API │ └───────┬──────────┘ ▼ ┌──────────────┐ │ LLM API │ └──────────────┘这种模式的好处无侵入业务代码不需要修改统一治理所有 LLM 调用走同一套监控逻辑可插拔可以叠加缓存、限流、熔断等中间件六、生产部署建议6.1 存储策略# 生产配置建议 storage: # 详细记录保留 7 天 detailed_records: retention: 7d storage: clickhouse # 聚合指标保留 90 天 aggregated_metrics: retention: 90d storage: victoria_metrics # 告警事件保留 180 天 alert_events: retention: 180d storage: elasticsearch6.2 采样策略sampling: # 错误/超时/截断100% 采样 error: 1.0 timeout: 1.0 truncated: 1.0 # 正常请求按模型分级采样 gpt-4o-mini: 0.05 # 低成本模型5% 采样 gpt-4o: 0.20 # 高成本模型20% 采样 claude-3-opus: 0.50 # 极高成本模型50% 采样6.3 告警阈值alerts: - name: P99 延迟 5s severity: warning interval: 5m - name: 单用户日成本 $100 severity: critical notify: finance-team - name: 模型错误率 10% severity: critical notify: on-call - name: 截断率 15% severity: warning action: auto-increase-max-tokens七、关键要点LLM 调用是最贵的环节​ — 必须精确追踪每一次调用的成本和延迟Client 中间件模式​ — 无侵入地给所有 LLM 调用加上监控按模型、按用户、按时段聚合​ — 三个维度缺一不可告警要分等级​ — 延迟高是 warning成本失控是 critical采样策略决定存储成本​ — 错误必采正常按比例成本优化是持续的过程​ — 监控数据驱动模型选择和 Prompt 优化 开发之余的小工具推荐监控 LLM 调用时经常需要计算 Token 消耗和成本。zz365.top 的数字转大写工具可以将美元金额转换为中文大写方便财务报销和合同填写。JWT 解析器可以快速解码用户认证 Token 中的 UserID 等信息。所有工具纯前端本地计算你的监控数据不会上传到服务器。下一讲预告​ 第4讲「决策质量监控」—— Jev 决策的观测、决策漂移检测、误判回捞、用户反馈闭环。
网站建设高端定制企业官网
RELATED

相关资讯

更多精彩内容,欢迎继续阅读

较早相关资讯

最新相关资讯

基于Swin Transformer的遥感变化检测实战:双时相影像判读 2026/10/1 22:54:30

基于Swin Transformer的遥感变化检测实战:双时相影像判读

/* MD / 富文本中的 .toc(含博客园搬家等嵌套结构);.toc-box 在侧栏,不受影响 */#content_views .toc,/* 编辑器常在目录前后插入空 p(:empty 仍占 20px),一并去掉避免顶空隙 */#content_views.markdown_views > p:empty:has(+ .toc),#content_views.markdown_views …

阅读更多 →
Manjaro/Arch下Fcitx5搜狗输入法实战安装指南 2026/10/1 22:54:16

Manjaro/Arch下Fcitx5搜狗输入法实战安装指南

1. 项目概述:为什么在 Manjaro/Arch Linux 上装搜狗输入法是个“高频痛点”Manjaro 和 Arch Linux 用户里,十有八九都卡在同一个地方:装完系统,打开浏览器想搜点东西,结果键盘敲出的全是英文字母——中文输入法没起来。…

阅读更多 →
AI生成代码如何匹配团队风格与工程一致性 2026/10/1 22:54:16

AI生成代码如何匹配团队风格与工程一致性

1. 这不是代码问题,是团队认知断层的显影 “AI写的代码一跑就通,但完全不像我们组写的”——这句话最近在好几个技术团队的茶水间、站会间隙、甚至代码评审会上反复出现。它听起来像一句调侃,但背后藏着一个正在快速扩大的现实裂口&#xff1…

阅读更多 →
示波器核心参数与实战技巧:带宽、采样率、探头和触发全解析 2026/10/1 22:54:16

示波器核心参数与实战技巧:带宽、采样率、探头和触发全解析

1. 示波器到底在“看”什么:从一堆波形说起我第一次拿起示波器的探头时,心里想的是“这不是个大号的万用表嘛”。后来被老师傅纠正了——万用表告诉你“现在是多少伏”,示波器告诉你“电压在过去这段时间里是怎么变的”。这一个“怎么变”的差…

阅读更多 →
ESXi虚拟机导出导入的底层逻辑与OVF/OVA交付实战 2026/10/1 22:54:16

ESXi虚拟机导出导入的底层逻辑与OVF/OVA交付实战

1. 为什么“导出导入虚拟机”不是点几下鼠标的事——ESXi环境下的真实交付瓶颈你有没有遇到过这样的场景:在Dell R730服务器上部署完ESXi 8.0,搭好vSphere环境,创建了一台CentOS 7 Hadoop 3.3 Spark 3.3伪分布式集群的虚拟机,测…

阅读更多 →
LabVIEW活用ActiveX生成Excel报表:不装NI报表工具包也能搞定 2026/10/1 22:54:16

LabVIEW活用ActiveX生成Excel报表:不装NI报表工具包也能搞定

如果我说,用LabVIEW生成Excel报表,不一定非要装NI Report Generation Toolkit,可能很多人第一反应是不信。毕竟网上90%的教程翻来覆去就是那套流程:拖出New Report.vi、Append Table To Report.vi,然后万事大吉。但这些…

阅读更多 →

今日资讯

本周资讯

本月资讯

看完文章仍有疑问?

联系尧图顾问,获取一对一建站咨询

立即免费咨询 📞 400-888-8888
📞 ✉