大模型可观测性实战:用Prometheus+Grafana搭建推理服务监控告警体系(TaoToken统一接入)

发布时间:2026/10/10 22:56:28
大模型可观测性实战:用Prometheus+Grafana搭建推理服务监控告警体系(TaoToken统一接入)
1. 推理服务上线后为什么必须补上可观测性大模型推理服务有个很反直觉的特点它很少直接崩掉更多时候是带病运行。进程还在、端口还通、健康检查也返回 200但首字延迟已经从 400ms 悄悄爬到 3 秒KV Cache 水位逼近 90%某个模型的回答质量开始下滑。这些问题用户往往能忍一两天才投诉等你收到反馈时体验分已经烧掉一大截。传统 Web 服务那套监控QPS、错误率、CPU、内存照搬到推理服务上会失灵。原因在于大模型推理有自己专属的生命体征首字延迟 TTFT 决定用户觉得卡不卡逐 token 延迟 TPOT 决定回答流得快不快KV Cache 使用率决定距离 OOM 还有多远排队请求数决定容量是否已经吃紧。这些指标在通用监控模板里根本没有必须单独设计。这篇面向已经部署了推理 API 的工程团队把整套可观测性体系搭起来指标怎么选、Prometheus 怎么抓、Grafana 看板怎么摆、告警规则怎么写、怎么用 TaoToken 统一 Key 和 API 通道接入后做端到端验证。全程给出可直接复制的配置片段跟着做就能跑通。适合谁看已经把 vLLM 或其他推理引擎推上生产、需要 7×24 保障的工程师正在做多模型统一接入、想给网关加监控的团队以及被服务到底稳不稳这个问题问住过的人。先说结论可观测性不是加几个 Panel 就完事它是一条从指标采集、看板呈现、告警触达到排障动线的完整链路。下面按这条链路一步步落地。2. TaoToken 统一接入与监控前置准备在动手配 Prometheus 之前先把接入层理清楚。很多团队的问题是推理服务有好几个实例、好几个模型调用方拿到的 Key 五花八门监控数据散落在各处根本没法按调用方维度做归因。这时候用 TaoToken 做统一接入层就很合适——它把模型调用收敛到一个 API 通道Key 统一管理指标也能按统一维度打点。TaoToken 在这里扮演的角色是统一 API 网关你的推理服务、外部模型、不同调用方都通过同一个 Base URL 和 Key 走。官网地址是 https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content API 端点是 https://taotoken.net/api 注意 API 地址不加 UTM 参数。前置准备分三步。第一步拿到统一 Key。进入控制台创建 API Key路径是 https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_contentconsoleutm_campaignrewrite Key 管理页在 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 。创建后立刻复制保存页面刷新后不再完整显示。第二步确认模型 ID。不同模型的 Model ID 不一样接入文档在 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 里面有完整的模型列表和参数说明。监控打点时按 Model ID 分面才能看出是哪个模型拖慢了整体延迟。第三步规划指标维度。统一接入后建议在网关层按model、caller、status三个标签打点。model区分模型caller区分调用方谁在用、用了多少status区分成功失败。这三个维度决定了你后面能不能做成本归因和故障定位。如果你用的是 Claude Code 这类编码工具接入配置三件套是 Base URL、Key、Model ID缺一不可。Base URL 填https://taotoken.net/apiKey 填刚创建的Model ID 按文档选。Coding Plan 适合长期编码和 Agent 场景入口在 https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 。前置准备做完你应该手里有一个可用的 Key、确认过的 Model ID、规划好的指标标签。接下来进入配置环节。3. Prometheus 抓取配置与 Grafana 看板 JSON这一节是核心给出可直接复制的配置。先看 Prometheus 抓取配置。vLLM 原生暴露 Prometheus 端点接入成本几乎为零。在prometheus.yml里加一段global: scrape_interval: 15s evaluation_interval: 15s rule_files: - /etc/prometheus/alert_rules.yml scrape_configs: - job_name: vllm scrape_interval: 15s metrics_path: /metrics static_configs: - targets: [gpu01:8000, gpu02:8000] labels: env: prod role: inference - job_name: gateway scrape_interval: 15s static_configs: - targets: [gateway:9100] labels: env: prod role: gateway - job_name: node static_configs: - targets: [gpu01:9100, gpu02:9100]vLLM 原生暴露的关键指标直接可用vllm:time_to_first_token_seconds是 TTFT 分布vllm:time_per_output_token_seconds是 TPOT 分布vllm:gpu_cache_usage_perc是 KV Cache 水位vllm:num_requests_running是正在处理的请求数vllm:num_requests_waiting是排队长度vllm:request_success_total是成功计数。网关层补充业务维度指标用 Python 埋点from prometheus_client import Counter, Histogram, start_http_server REQ_COUNT Counter( gateway_requests_total, 请求数, [model, caller, status] ) REQ_LATENCY Histogram( gateway_latency_seconds, 端到端延迟, [model], buckets[0.1, 0.5, 1, 2, 5, 10, 30] ) TOKEN_COUNT Counter( gateway_tokens_total, Token消耗, [model, caller] ) def handle_request(req, caller): try: REQ_COUNT.labels(modelreq.model, callercaller[owner], statusok).inc() TOKEN_COUNT.labels(modelreq.model, callercaller[owner]).inc(req.tokens) except Exception: REQ_COUNT.labels(modelreq.model, callercaller[owner], statuserror).inc() raise start_http_server(9100)分工原则很清晰引擎指标看机器健不健康走 vLLM 端点业务指标看服务好不好用走网关埋点。两层都要有缺一层就是半个瞎子。Grafana 看板推荐四象限布局对应坏了→慢了→满了→多了四种故障模式。看板 JSON 的核心 Panel 配置如下{ panels: [ { title: TTFT P50/P99, type: timeseries, targets: [ { expr: histogram_quantile(0.50, sum(rate(vllm:time_to_first_token_seconds_bucket[5m])) by (le)), legendFormat: P50 }, { expr: histogram_quantile(0.99, sum(rate(vllm:time_to_first_token_seconds_bucket[5m])) by (le)), legendFormat: P99 } ] }, { title: KV Cache 水位, type: gauge, targets: [ { expr: vllm:gpu_cache_usage_perc, legendFormat: {{instance}} } ], fieldConfig: { defaults: { thresholds: { steps: [ {color: green, value: null}, {color: yellow, value: 0.7}, {color: red, value: 0.85} ] } } } }, { title: QPS 按模型分面, type: timeseries, targets: [ { expr: sum(rate(gateway_requests_total[5m])) by (model), legendFormat: {{model}} } ] }, { title: 错误率, type: timeseries, targets: [ { expr: sum(rate(gateway_requests_total{status\error\}[5m])) / sum(rate(gateway_requests_total[5m])), legendFormat: error_rate } ] } ] }排障动线先看错误象限有没有报错再看延迟象限慢不慢慢了看容量象限满没满最后看流量象限是否异常。按这个顺序扫一遍30 秒内完成初判。告警规则 YAML 六条保命线groups: - name: llm_service rules: - alert: HighTTFT expr: histogram_quantile(0.99, sum(rate(vllm:time_to_first_token_seconds_bucket[5m])) by (le)) 2 for: 5m labels: severity: warning annotations: summary: TTFT P99 超过2秒用户体验受损 - alert: KVCacheHigh expr: vllm:gpu_cache_usage_perc 0.8 for: 5m labels: severity: critical annotations: summary: KV Cache水位超80%逼近OOM - alert: QueueBacklog expr: vllm:num_requests_waiting 10 for: 10m labels: severity: warning annotations: summary: 排队请求持续超过10个考虑扩容 - alert: ErrorRateHigh expr: sum(rate(gateway_requests_total{statuserror}[5m])) / sum(rate(gateway_requests_total[5m])) 0.01 for: 5m labels: severity: critical annotations: summary: 错误率超1%引擎或网关异常 - alert: ThroughputDrop expr: rate(vllm:generation_tokens_total[10m]) 0.5 * rate(vllm:generation_tokens_total[1h] offset 1d) for: 10m labels: severity: warning annotations: summary: Token生成速度相比昨日同期腰斩 - alert: InstanceDown expr: up{jobvllm} 0 for: 1m labels: severity: critical annotations: summary: vLLM实例失联告警纪律两条告警必须可行动收到告警不知道该怎么办的不如不配分级触达水位类发群消息、实例失联打电话别让半夜的电话只为一条水位75%响起。4. 用 curl 验证指标端点与 PromQL 核对 QPS 延迟配置写完必须验证否则你永远不知道数据有没有真的进来。这一节给出完整的验证动作。先验证 vLLM 指标端点是否正常暴露curl -s http://gpu01:8000/metrics | grep -E vllm:(time_to_first_token|gpu_cache_usage|num_requests)正常输出应该能看到类似vllm:time_to_first_token_seconds_bucket{le0.5} 120 vllm:time_to_first_token_seconds_bucket{le1.0} 340 vllm:gpu_cache_usage_perc 0.42 vllm:num_requests_running 8 vllm:num_requests_waiting 2如果这条命令返回空说明端点没挂上或者路径不对先排查引擎启动参数。再验证 Prometheus 是否抓到了目标curl -s http://prometheus:9090/api/v1/targets | jq .data.activeTargets[] | {job: .labels.job, health: .health}health应该是up。如果是down看lastError字段通常是网络不通或端口写错。接着用 PromQL 核对 QPS 和延迟曲线。在 Prometheus 查询界面或 API 里执行# 核对 QPS curl -s http://prometheus:9090/api/v1/query?querysum(rate(gateway_requests_total[5m])) | jq # 核对 TTFT P99 curl -s http://prometheus:9090/api/v1/query?queryhistogram_quantile(0.99,sum(rate(vllm:time_to_first_token_seconds_bucket[5m]))by(le)) | jq # 核对 KV Cache 水位 curl -s http://prometheus:9090/api/v1/query?queryvllm:gpu_cache_usage_perc | jq如果 QPS 返回 0 但服务明明在跑检查网关埋点的start_http_server端口是否和 Prometheus 抓取目标一致。如果 TTFT 返回 NaN通常是 histogram bucket 还没数据等一个抓取周期再看。最后做一次端到端验证通过 TaoToken 统一通道发一个请求然后立刻查指标有没有变化。curl -s https://taotoken.net/api/v1/chat/completions \ -H Authorization: Bearer $TAOTOKEN_KEY \ -H Content-Type: application/json \ -d { model: your-model-id, messages: [{role: user, content: ping}], max_tokens: 16 }请求成功后回到 Prometheus 查gateway_requests_total应该 1vllm:num_requests_running在请求期间有波动。这一步跑通说明从接入层到监控层的链路完整了。模型对话功能可以在 https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_contentchatutm_campaignrewrite 直接验证适合快速确认 Key 和 Model ID 是否配对正确。5. 常见报错排查401、local proxy failed、reading choices、OAuth这一节对照真实报错给出排查路径。这些坑我基本都踩过。401 Unauthorized。最常见的原因是 Key 没带对或者带了多余空格。检查Authorization: Bearer $TAOTOKEN_KEY里 Key 是否完整、有没有换行符。另一个原因是 Key 被禁用或过期去 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 确认状态。如果用的是 Claude Code 接入检查 settings 里的 Base URL 是不是https://taotoken.net/api末尾不要多加斜杠。local proxy failed。这个报错通常出现在本地代理配置环节指的是请求没能到达目标端点。排查顺序先确认 Base URL 拼写正确再确认网络能通curl -v https://taotoken.net/api最后检查本地环境变量有没有残留的代理设置干扰。注意不要配置任何非官方的转发通道直接用官方 API 地址即可。reading choices 相关报错。这类报错一般出现在解析响应体时choices字段读不到。原因可能是返回的不是标准 chat completions 格式比如请求被路由到了错误的端点或者 Model ID 写错导致返回了错误结构。检查请求路径是不是/v1/chat/completionsModel ID 是否和文档一致。接入文档在 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 对照确认。OAuth 相关报错。如果用的是 Claude Code 或类似工具OAuth 报错通常是认证流程没走完或者 token 过期。Claude Code 的接入配置在 https://taotoken.net/claude-code?utm_sourcetaotoken_aicg_blog_endutm_contentclaudecodeutm_campaignrewrite 按文档重新走一遍认证。注意 Base URL、Key、Model ID 三件套要同时正确缺一个都会报认证失败。指标端点返回 404。vLLM 的/metrics路径如果 404检查启动参数里有没有开启 metrics。部分版本需要显式指定--enable-metrics或类似参数。另外确认metrics_path配置和实际路径一致。Prometheus target down。先curl目标端点确认服务活着再检查防火墙和端口。如果是容器环境注意容器网络和宿主机网络的差异localhost在容器里指向容器自己。Grafana 看板无数据。检查数据源配置的 Prometheus 地址是否正确以及查询时间范围是否覆盖了数据。如果 PromQL 在 Prometheus 里能查到但 Grafana 查不到多半是数据源 URL 写成了localhost而 Grafana 在另一个容器里。排查的核心思路先确认单点能通curl 端点再确认采集能通Prometheus target最后确认展示能通Grafana 查询。逐层排除不要跳步。6. 把监控接入长期编码与 Agent 工作流监控体系搭好之后下一步是让它服务于长期运行的工作流。如果你在做 Coding Agent 或者长期编码任务推理服务的稳定性直接决定任务能不能跑完。这时候 Coding Plan 配合监控看板就很实用入口在 https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 。具体做法把告警规则里的caller维度细化到具体任务或 Agent 实例这样某个 Agent 疯狂重试导致 QPS 突增时你能立刻定位到是哪个任务在刷。成本 Panel 按caller聚合 Token 消耗某调用方用量突增 300%先问是不是脚本写死了循环。质量观测也别忘。技术指标保证服务活着质量观测保证服务聪明。每天随机抽 1% 的问答对入库人工或 LLM-as-Judge 抽检模型版本升级后必做对比抽检。指标全绿但回答变蠢的情况只有这个能发现。最后给一个实用技巧把排障动线做成 Grafana 的快捷链接告警消息里直接带上跳转 URL。收到告警点一下就到对应看板省去手动切时间范围和筛选条件的时间。告警分级也要落实水位类发群、实例失联打电话别让所有告警都走同一个通道。整套体系跑通后你对推理服务的状态就从祈祷不出事变成了知道自己在哪。指标端点、看板、告警、排障动线四件套齐了服务才算真正上线。