Agentic Work Atlas

标签: benchmark

此标签下有9条笔记。

  • 2026年7月27日

    Frontier Models with Our Harness Achieve ~99% on ARC-AGI-3 Public — Schema

    • source-summary
    • agent-harness
    • benchmark
    • AGI
    • agentic-engineering
  • 2026年7月18日

    20260717-schema-harness-arc-agi

    • clippings
    • agentic-engineering
    • agent-harness
    • AI-Agent
    • benchmark
  • 2026年7月08日

    Minimal Pair Evaluation

    • evaluation-methodology
    • agentic-engineering
    • benchmark
  • 2026年6月18日

    20260617-openai-lifescibench

    • clippings
    • benchmark
    • life-science
    • evaluation
    • frontier-model
    • GPT-Rosalind
  • 2026年6月18日

    OpenAI LifeSciBench: 750 任务生命科学 Benchmark

    • source-summary
    • benchmark
    • evaluation
  • 2026年6月16日

    CORE-Bench

    • benchmark
    • reasoning
    • AI-evaluation
  • 2026年6月16日

    Einstein Test

    • AI
    • benchmark
  • 2026年6月16日

    ITBench

    • agent-evaluation
    • agentic-engineering
    • benchmark
  • 2026年6月16日

    SWE-Bench

    • benchmark
    • software-engineering
    • AI-evaluation

Created with Quartz v4.5.2 © 2026

  • GitHub