1 Matching Annotations
  1. Last 7 days
    1. This is especially true for public benchmarks, which model makers can easily hill climb by simply creating a bunch of RL environments that closely mimic the benchmark tasks.

      作者给自己结论留的后门,也是最该记住的一条:公开基准可以用定制 RL 环境直接刷分。于是"开源已追上"的证据强度取决于评测是否私有,而本文所用基准恰恰多为公开——结论方向可信,幅度存疑。