Agentic Benchmark Checklist

编译摘要

1. Benchmark validity precedes model scoring

ABC treats the benchmark as a measurement system that must itself be audited:

task validity
  ×
outcome validity
  ×
reporting / provenance

A high-quality grader cannot rescue a task whose allowed solutions, environment or reference are wrong.

2. Semantic equivalence is part of reference integrity

A reference should not encode one implementation path as the only truth when multiple outputs are functionally valid.

This makes reference truth broader than a static answer key.

3. Oracle/human checks are construction controls

Oracle solvers, human baselines and judge-human agreement are useful checks, but the framework does not imply that every evaluated run has an external oracle.

4. 对标

关联概念