5 Matching Annotations
  1. Sep 2026
    1. evaluations like Petri are only proxies for real-world misalignment, and we did not test whether alignment gains persist after extensive RL training on other tasks.

      【局限】研究人员明确指出,Petri等评估工具仅是现实世界中错位行为的代理指标,而非完全等效的替代。此外,研究未验证对齐改进在其他任务上进行大量RL训练后是否能保持稳定,这限制了结果的实用价值。

  2. Aug 2026
    1. A physical evaluation tests an AI system in the actual physical world — not a simulator, not a sandbox, not a virtual environment dressed up as one.

      Physical evals vs simulations

      Physical evals measure AI performance in real environments, avoiding simulation edge cases and capturing real-world complexity that virtual systems can't replicate.

  3. May 2026
    1. AI Village gives multiple AI agents their own computer environments and a shared group chat, then tasks them with open-ended real-world goals like fundraising, organizing events, making games, and gaining subscribers.

      这个案例展示了开放世界评估的实际应用,每年约5万美元的成本表明这种评估需要相当大的资源投入。相比传统基准测试,这种评估方式更接近真实应用场景,但也因此成本更高,难以大规模实施。

  4. Nov 2021
  5. Jul 2020