OpenAI SWE-bench Verified audit
编译摘要
1. Reference/test quality can dominate the remaining failures
OpenAI's audit found material prompt/test issues in 59.4% of the selected 138-task subset that o3 failed to solve consistently over 64 runs.
The relevant unit is therefore:
selected hard-task subset
→ expert specification/test audit
→ benchmark-quality judgmentnot the full 500-task benchmark.
2. Independent human review improves benchmark construction, not run-level truth
At least six engineers independently reviewed each selected task, with additional re-verification for flagged issues.
This is strong evidence that benchmark reference/test quality itself requires an oracle-audit layer. It is not evidence that every model run received an independent external verdict.
3. Contamination and reference validity are different failure modes
The same post reports models reproducing gold-patch/problem details for some tasks.
Therefore:
bad tests / bad specification
≠
training contaminationBoth can corrupt a score, but by different causal paths.
4. 对标
- 对 Verifiable-Agent-Engineering:reference truth 必须与 execution truth 分开。
- 对 Evaluation-Integrity:benchmark contamination and benchmark validity are distinct integrity dimensions.