τ³ task-reference fixes

编译摘要

1. Wrong expected actions can make correct behavior score as failure

The official audit documents tasks whose expected actions contradicted the domain policy or available tools.

This is direct benchmark-maintenance evidence that a false negative can originate in the reference contract, not the agent.

2. Task validity and reference validity interact

Other fixes changed ambiguity, impossible constraints, fallback behavior and policy loopholes.

Therefore the before/after score shift cannot be attributed only to a corrected answer key.

3. Versioning matters

A benchmark score without task/evaluator version can become uninterpretable after task definitions change.

The minimum comparison identity should therefore bind:

benchmark version
  + task id
  + policy / environment version
  + evaluator / expected-action version
  + model / trial

4. 对标

关联概念