OpenAI SWE-bench Verified audit

编译摘要

1. Reference/test quality can dominate the remaining failures

OpenAI's audit found material prompt/test issues in 59.4% of the selected 138-task subset that o3 failed to solve consistently over 64 runs.

The relevant unit is therefore:

selected hard-task subset
  → expert specification/test audit
  → benchmark-quality judgment

not the full 500-task benchmark.

2. Independent human review improves benchmark construction, not run-level truth

At least six engineers independently reviewed each selected task, with additional re-verification for flagged issues.

This is strong evidence that benchmark reference/test quality itself requires an oracle-audit layer. It is not evidence that every model run received an independent external verdict.

3. Contamination and reference validity are different failure modes

The same post reports models reproducing gold-patch/problem details for some tasks.

Therefore:

bad tests / bad specification
        ≠
training contamination

Both can corrupt a score, but by different causal paths.

4. 对标

关联概念