Claude on call: How Claude Tag serves as Anthropic's first responder for CI/CD failures
编译摘要
1. 浓缩
核心结论1: Anthropic 用 Claude Tag 把 CI/CD on-call 的 first responder 完全代理给 agent——首份 SITREP 15 分钟内、最快 4 分钟定位根因、revert 后 3 分钟内验证告警恢复
关键证据: median 14 minutes first evidence-grounded analysis; fastest 4 minutes; 3-minute post-revert Slack ping; "Claude authored the first situation report in every recent incident that had one"
关键证据: "orchestration agent spins up executor subagents to investigate each dependency and source of truth"; "Claude can chase multiple leads in parallel, helping to reduce MTTR"; 617-line investigation skill for shadow divergence bugs
关键证据: "Claude appends to it on its own automatically"; "If the same pattern shows up enough times, we promote it into the investigation skill itself"; 经典条目 "query the data first, then theorize"
2. 质疑
关于"15 分钟首份 SITREP"的可迁移性: 这是 Anthropic CI 团队自身基础设施(Datadog/Grafana/GitHub/K8s + Claude Tag service account)下的指标,外部组织需要先打通这些 MCP connector。文中给出的 on-call-kit 模板可压缩搭建时间,但 connector 配置的真实成本未量化
跨域关联3: on-call first responder 是 Operational-Responsibility "first person paged = 最佳修复者"的延伸——AI 接管"被 paging"的部分,但责任链仍归人("Either of us can steer the investigation or add a hypothesis in real-time, together")。这是 Palantir 预言的 agent 责任路由问题的实证答案