9 Matching Annotations
  1. Aug 2026
    1. when Claude Opus 5 was run at Max reasoning effort, it scores 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2 semi-private

      被忽略的第三个自由变量:推理预算。30.2% 那条基线是 high effort,而 AVO 的跑法用了另一套 reasoning setting。同一模型换档位分数就能大幅移动,所以任何跨系统比分表,都得先标注 effort 档位、观测格式和评测集,再谈差值。

    1. In the AVO configuration, the LLM operated in a text-only modality: each observation was supplied as an exact 64 x 64 text grid, with no images or image tokens sent to the model.

      与 VISTA 对照时的关键混杂项:对方主配置喂 512×512 渲染图,AVO 直接喂 64×64 文本网格。感知通道都换了,动作效率自然不可比。要警惕把“观测表示的胜利”读成“智能体架构的胜利”——真正的运行单元是模型×harness×工具环境×上下文策略。

  2. Sep 2020
  3. Aug 2020
  4. Jul 2020