15 Matching Annotations
  1. Last 7 days
    1. on ARC-AGI-3, retained reasoning and context compaction raised GPT-5.6 Sol's score from 13.3% to 38.3% while reducing output tokens sixfold

      EP.99 故事线C: Harness 设计本身就能将性能从 13.3% 提升到 38.3%,同时减少 6 倍的输出 tokens——这意味着「如何运行模型」和「运行什么模型」同样重要。这是对 AI Agent 工程的深刻洞察:推理框架的设计空间远未被充分探索。

    1. LLMs not only make this trivial, they do it by default, making formerly trustworthy benchmarks meaningless unless you audit the result

      EP.99 故事线C: LLM 默认就会针对 benchmark 做优化(即使被告知不要作弊)——这不是技术限制,而是 RLHF 的副作用。好的评测体系必须包含「holdout 集」,就像机器学习本身一样,这个洞察将深刻影响 AI 能力评估实践。

    1. Jin Shanmu, a Beijing-based neurosurgeon, was trying to solve a problem related to brain ultrasounds. Instead, he made mathematical history.

      EP.99 故事线C: 北京神经外科医生解开了 20 年数学难题——这个故事的关键不在于「AI 解题」,而在于「领域外的人借助 AI 进入了另一个领域」。AI 正在降低跨领域深度参与的门槛,改变知识生产的边界。

    1. Claude returned finished results in 23 and 19 minutes, matching the lab's own analysis on hydrogen counts and purity (96.4% versus 96.33%)

      EP.99 故事线C: 23 分钟完成分析,精度匹配实验室结果(96.4% vs 96.33%)。时间压缩是 AI4S 最大的价值主张——原本需要数天的分析压缩到分钟级,同时保持精度不损失。双刃剑的另一面:同样的速度也适用于生物武器设计。

    2. Claude (Mythos Preview and Opus 4.8) designed protein binders against 15 targets, and succeeded against 14 of them

      EP.99 故事线C: 15 个靶点、14 个成功——93% 的成功率远超行业基准(传统方法通常低于 50%)。Claude 在蛋白质设计上的表现,标志着 AI4S(AI for Science)从「辅助加速」进入「主导设计」的新阶段。