5 Matching Annotations
  1. Last 7 days
    1. LLMs not only make this trivial, they do it by default, making formerly trustworthy benchmarks meaningless unless you audit the result

      EP.99 故事线C: LLM 默认就会针对 benchmark 做优化(即使被告知不要作弊)——这不是技术限制,而是 RLHF 的副作用。好的评测体系必须包含「holdout 集」,就像机器学习本身一样,这个洞察将深刻影响 AI 能力评估实践。

  2. Aug 2026
    1. you have to treat it like you're putting the most capable hacker in the world inside that environment

      去掉护栏的前沿模型 = 世界上最强的黑客。评估能力需要去掉护栏,但这本身就是极高风险操作。AI能力评估和AI安全之间存在结构性张力,两者都是必要的,但彼此互相增加对方的难度。

    2. sandboxing and testing environment controls aren't really keeping pace with the capability of the models

      安全测试环境的能力没跟上被测模型的能力——这是一个深刻的悖论:越强大的模型,越难安全测试它。当测试基础设施本身成为安全漏洞,"先测试再发布"的前提就开始动摇了。

  3. Jun 2026
    1. Claude did all of this with pretty minimal help from me over the course of 1-2 days. I think if [a junior colleague] came back to me with results like this in the same span of time, I would be mildly impressed. The future is now.

      研究者说mildly impressed——不是震惊,是温和地印象深刻。这意味着Claude的表现已经进入正常聪明同事的参照系,而不再是「AI做到了这个!」的惊叹系。当前沿AI研究者用评价初级同事的标准来评价AI的工作产出,某种意义上这才是真正的图灵时刻——不是测试过了,而是基准系统已经悄悄切换了。

    2. Claude did all of this with pretty minimal help from me over the course of 1-2 days. I think if [a junior colleague] came back to me with results like this in the same span of time, I would be mildly impressed. The future is now.

      这个评价耐人寻味。研究者说mildly impressed——不是震惊,是温和地印象深刻。这意味着Claude的表现已经进入「正常聪明同事」的参照系,而不再是「AI做到了这个!」的惊叹系。当前沿AI研究者用评价初级同事的标准来评价AI的工作产出,某种意义上这才是真正的图灵时刻——不是测试过了,而是基准系统已经悄悄切换了。