6 Matching Annotations
  1. Last 7 days
    1. Adding only retained reasoning and compaction, GPT-5.6 Sol’s ARC-AGI-3 score tripled from 13.3% to 38.3%.

      比 AVO 的满分更有说服力:只动“保留推理”和“上下文压缩”两个开关,基线同源、变量可数,这才是消融实验该长的样子。但同样只在公开集上,且起点 13.3% 很低——低分区的三倍不能线性外推到高分区,压缩带来的收益通常随基线上升而衰减。

    1. Nvidia's AVO agent system lifted Claude Opus 5 from a 30% baseline to 100% on the ARC-AGI-3 reasoning benchmark, across all 183 levels.

      把 30% 和 100% 相减当成 harness 的贡献,是这轮报道里最常见的读法错误。两个数字来自不同评测设置:一边是 ARC Prize 跑裸模型,一边是 NVIDIA 自建观测接口、记忆与监督器的完整系统。NVIDIA 原文明说过这不是受控消融,本文没有转述这句限定。

    1. These results cover the 25-environment ARC-AGI-3 public set using the official scorecard and RHAE metric. They are not results on the semi-private or fully private competition sets.

      满分只在公开集上成立。ARC 设半私有/私有集,正是为了拦住针对公开题目的过拟合与反复调参,因此公开集 100 分不能推出私有集能力。任何把它转述成“达到人类水平”的说法,都已经丢掉了这句限定。

    2. This should not be interpreted as a controlled ablation: the two systems differ in agent backend, observation representation, memory, context management, and other implementation details.

      全文最该被引用的一句:NVIDIA 自己先声明这不是受控消融。也就是说 12% 的动作数优势里,后端模型、观测表示、记忆、上下文管理四个变量同时在变,无法归因给任何单一设计。读跑分新闻时,这类作者自认的免责声明信息量往往大于标题数字。

  2. May 2026
    1. ARC-AGI-3 was officially released this week. All frontier models score below 0.5%

      ⚠️【令人震惊的数字】最强前沿模型得分低于 0.5%——而非专业人类轻松超过 60%,差距超过 120 倍。这是继 ARC-AGI-2 之后最彻底的「AI 能力幻觉清醒剂」。推理能力的提升并未自动迁移到「新颖抽象推理」,当所有人在讨论 AGI 即将到来时,这份数据是最直接的反驳。