ELT-Bench-Verified

编译摘要

1. Evaluator/reference correction can change measured capability while agent output is held fixed

The audit isolates benchmark-side defects including rigid comparison logic, ambiguous specifications and incorrect ground truth.

After correcting evaluation logic and removing unreliable ground-truth columns, the same SWE-Agent + Claude Sonnet 4.5 setup moves from 22.66% to 32.51% transformation success.

This is direct evidence that:

measured capability
  = f(agent output, reference, evaluator)

not agent output alone.

2. Human disagreement can imply that no authoritative replacement exists

For 30 suspect GT columns, three independent engineers achieved only 57.8% average pairwise exact-match agreement.

The correct response was removal, not majority-vote fabrication of a new answer key.

3. Semantic equivalence belongs in the evaluator contract

The verified evaluator normalizes booleans, floating-point tolerance, percentage/decimal formats, NULL representation and row order where order is unspecified.

Therefore a strict string/value representation match can create false failures even when transformation semantics are acceptable.

4. Corrected benchmark scores need version identity

A result should bind at least:

task/spec version
  + GT/reference version
  + evaluator normalization rules
  + agent/model version

Otherwise pre/post-correction results can be mistaken for model progress or regression.

5. 对标

关联概念