关键证据: "Memory keeps paying off as long as a model has a remaining failure mode to target"; SGC metric gains exceed TGC because guidelines help clear all variants of a scenario
核心结论3: Prompt caching 是 production 经济杠杆——DeepSeek 在 full guideline set 下 tokens/task 从 148K 涨到 263K(+78%),但若保持 guideline-set prefix 稳定可被 cache;同时 curated retrieval 在 weak model 上既最准又最便宜(gpt-oss-120b +16.1pp at +5% tokens)
关于 "saturated 模式" 的因果:作者诚实承认 "the label describes what we observed, not a proven cause"——GLM-5 零增益可能是任务已接近模型天花板、guideline 未覆盖剩余失败、或模型未有效应用 guidance,无法从单次 sweep 区分。本研究未做 controlled ablations 隔离这些因素
关于 AppWorld 的代表性:单一 benchmark(585 tasks, 9 simulated apps)虽是"rigorous multi-step benchmark, but a single one"——calendar/messaging/payments 等结构化任务是否能代表开放域 agent 工作流未验证
关于 "context window size is hypothesized but not isolated":作者明示这是假设,未做 controlled experiments 隔离 context window 与 raw capability。这让"saturated vs selective"的预测因子清单(headroom / context window / architecture / guideline quality / task distribution)尚处于观察性归纳
关于 curated retrieval 的 ranking 质量:cosine similarity ranking 不完美预测哪些 guideline 帮助给定任务——means selective 配置仍可被更好 selector 改进;这是 "What's next" 的核心 open problem
关于 "no human annotation" 的覆盖:ALTK-Evolve 不需要人类标注,但需要 mining 阶段的 trajectory 足够覆盖失败模式;若 agent 极少失败某类问题,guideline 集中不会包含该类 fix
3. 对标与旁逸
跨域关联1: guideline set mining from past trajectories 与 Anthropic Claude Tag 的 lessons.md→investigation skill 自蒸馏循环同构——Anthropic 是定性叙述,IBM 是定量 8 模型 sweep。两者合并给出 "agentic memory 不仅是蒸馏,更是剂量调节" 的双向证据
跨域关联2: 三种剂量模式 与 Structured-Agent-Memory 的属性匹配范式互补——后者讲记忆怎么组织,前者讲该喂多少。可组合:结构化记忆按 dosage 选择 full set vs retrieval