第 2 课
只读排查:列应用、查错误事故与性能
按官方 Tool Reference:get_applications → get_exception_incidents → get_incident → get_performance(可选 get_traces / get_log_lines)。
学习位置仅保存在当前浏览器,有效期 180 天。
图文讲义
来源:AppSignal · MCP Tool Reference、AppSignal MCP。 本课假设你已完成第一课鉴权;工具名以官方 Reference 为准。优先做只读调用;写操作(
update_/manage_)会改生产配置,需单独在客户端里允许。
你将得到什么
- 会触发官方只读工具:
get_applications、get_exception_incidents、get_incident、get_performance(及可选get_traces/get_log_lines) - 完成「先列应用 → 查开放错误 → 看慢动作」的闭环
- 知道多数工具需要
app_name+app_environment;名称大小写不敏感
官方工具一览(本课用到的)
| 工具 | 读写 | 作用 |
|---|---|---|
get_applications | 读 | 列出可访问的 app_name/app_environment |
get_exception_incidents | 读 | 搜索/分页列出错误事故(默认每页 50) |
get_incident | 读 | 按事故编号取详情、堆栈、状态 |
get_performance | 读 | 性能概览:慢动作 / 采样 traces |
get_traces | 读 | 按 action 或 digest 拉 trace 树与 span |
get_log_lines | 读 | 用 AppSignal 表达式查日志 |
自然语言描述任务即可;Agent 负责选工具填参。需要纠正参数时,直接说出官方参数名(如 states: open、revision: last)。
步骤 1:列出应用
发送:
Use get_applications and list every app_name/app_environment I can access.
对照结果:出现应用列表。记下你要排查的一对,下文用 MyApp / production 作占位——请换成你列表里的真实值。
步骤 2:列出开放错误事故
发送(把应用名换成你的):
For MyApp production, use get_exception_incidents with states open and show the top incidents: number, exception name, action, last occurrence.
对照结果:一页开放事故(最多约 50 条)。记下一条感兴趣的 incident number(数字 ID)。
可选收窄(官方参数):
query:对异常名/动作名做关键词搜索(支持*)revision: "last":只看最近一次部署引入的错误namespaces:如web,background
步骤 3:打开单条事故详情
发送:
Use get_incident on MyApp production for incident number <NUMBER>. Summarize state, severity, first/last seen, and the top of the stack trace.
对照结果:状态(open / closed / wip)、指派、堆栈摘要。先不要调用 update_incidents 改状态,除非你明确要在客户端批准写权限。
步骤 4:查看性能概览
发送:
For MyApp production, use get_performance and list the slowest actions with mean duration and error rate if available.
对照结果:慢动作或采样性能 traces(Ruby/Elixir 采样;若应用上报 OpenTelemetry,官方还会给出按动作排序的列表)。挑一个慢 action_name(如 BlogPostsController#index)进入可选加深:
Use get_traces for MyApp production in namespace web with action_name BlogPostsController#index, list recent slow traces, then open the slowest trace_id as a span tree.
对照结果:先 list 得到 trace_id,再 tree 模式看到 span 树与按类别的 self time 分解。
步骤 5:(可选)查错误相关日志
若该应用已接入 Logging:
For MyApp production, use get_log_lines with query severity=error and limit 20. Summarize the newest messages.
官方表达式示例:severity=error、message:timeout、severity=error AND hostname=web-1。空 query 表示拉取全部(仍受 limit 限制)。
安全与范围提醒
| 点 | 官方要点 |
|---|---|
| 读写分界 | get_ / discover_ 只读;create_ / update_ / manage_ 会改数据 |
| Token 收窄 | MCP Token 可按应用与工具集勾选;入门 Agent 可只开 discovery + exceptions |
| 无数据 | MCP 只是网关:应用没上报日志时 get_log_lines 为空属正常 |
| 反馈 | 缺能力时可让 Agent 调 get_more_tools 记一笔需求(不会解锁隐藏工具) |
本课小结
| 检查项 | 期望 |
|---|---|
get_applications | 有 app/env 列表 |
get_exception_incidents | 能列出 open 事故与 number |
get_incident | 能看到堆栈摘要 |
get_performance | 能看到慢动作或 traces |
完成以上四步,即达成「Agent 直连 AppSignal 做只读排查」的最小闭环。更多工具(anomaly、check-ins、dashboards、metrics)见官方 Tool Reference。
