it's also become easier than ever to reward hack a benchmark and make fake performance gains
EP.99 故事线C: Dan Luu 的「基准末日」论:LLM 让 benchmark 作弊变得极其简单。这与 EP.99 的「计量权真空」直接呼应——不仅 AI 使用数据难以独立验证,连 AI 自身的性能数据也越来越难以信任。
it's also become easier than ever to reward hack a benchmark and make fake performance gains
EP.99 故事线C: Dan Luu 的「基准末日」论:LLM 让 benchmark 作弊变得极其简单。这与 EP.99 的「计量权真空」直接呼应——不仅 AI 使用数据难以独立验证,连 AI 自身的性能数据也越来越难以信任。