Agentic Work Atlas
Search
搜索
暗色模式
亮色模式
Explorer
标签: benchmark
此标签下有9条笔记。
2026年7月27日
Frontier Models with Our Harness Achieve ~99% on ARC-AGI-3 Public — Schema
source-summary
agent-harness
benchmark
AGI
agentic-engineering
2026年7月18日
20260717-schema-harness-arc-agi
clippings
agentic-engineering
agent-harness
AI-Agent
benchmark
2026年7月08日
Minimal Pair Evaluation
evaluation-methodology
agentic-engineering
benchmark
2026年6月18日
20260617-openai-lifescibench
clippings
benchmark
life-science
evaluation
frontier-model
GPT-Rosalind
2026年6月18日
OpenAI LifeSciBench: 750 任务生命科学 Benchmark
source-summary
benchmark
evaluation
2026年6月16日
CORE-Bench
benchmark
reasoning
AI-evaluation
2026年6月16日
Einstein Test
AI
benchmark
2026年6月16日
ITBench
agent-evaluation
agentic-engineering
benchmark
2026年6月16日
SWE-Bench
benchmark
software-engineering
AI-evaluation