3 Matching Annotations
  1. Last 7 days
    1. pluralism is most decisively made or unmade at the deployment-governance layer: interfaces, preference-data pipelines, and audit infrastructure.

      This argument shifts the locus of the problem from the model's architecture to the socio-technical systems that surround it. It's a provocative claim that the core issue isn't 'how to build a better model' but 'how to build a better system for deploying and governing models,' placing the onus on developers and regulators, not just AI researchers.

  2. Apr 2026
    1. We built an automated scanning agent that systematically audited eight among the most prominent AI agent benchmarks — SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, and CAR-bench — and discovered that every single one can be exploited to achieve near-perfect scores without solving a single task.

      令人惊讶的是:研究人员构建的自动化扫描工具发现,所有八个主流AI代理基准测试都存在漏洞,无需解决任何任务就能获得接近完美的分数。这表明整个AI评估领域存在系统性问题,几乎所有当前使用的基准测试都不可靠。

  3. Oct 2020