ELT-Bench-Verified
编译摘要
1. Evaluator/reference correction can change measured capability while agent output is held fixed
The audit isolates benchmark-side defects including rigid comparison logic, ambiguous specifications and incorrect ground truth.
After correcting evaluation logic and removing unreliable ground-truth columns, the same SWE-Agent + Claude Sonnet 4.5 setup moves from 22.66% to 32.51% transformation success.
This is direct evidence that:
measured capability
= f(agent output, reference, evaluator)not agent output alone.
2. Human disagreement can imply that no authoritative replacement exists
For 30 suspect GT columns, three independent engineers achieved only 57.8% average pairwise exact-match agreement.
The correct response was removal, not majority-vote fabrication of a new answer key.
3. Semantic equivalence belongs in the evaluator contract
The verified evaluator normalizes booleans, floating-point tolerance, percentage/decimal formats, NULL representation and row order where order is unspecified.
Therefore a strict string/value representation match can create false failures even when transformation semantics are acceptable.
4. Corrected benchmark scores need version identity
A result should bind at least:
task/spec version
+ GT/reference version
+ evaluator normalization rules
+ agent/model versionOtherwise pre/post-correction results can be mistaken for model progress or regression.
5. 对标
- 对 Verifiable-Agent-Engineering:directly strengthens versioned reference/evaluator contract.
- 对 Evaluation-Integrity:shows benchmark correction alone can materially change reported capability.