when Claude Opus 5 was run at Max reasoning effort, it scores 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2 semi-private
被忽略的第三个自由变量:推理预算。30.2% 那条基线是 high effort,而 AVO 的跑法用了另一套 reasoning setting。同一模型换档位分数就能大幅移动,所以任何跨系统比分表,都得先标注 effort 档位、观测格式和评测集,再谈差值。
when Claude Opus 5 was run at Max reasoning effort, it scores 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2 semi-private
被忽略的第三个自由变量:推理预算。30.2% 那条基线是 high effort,而 AVO 的跑法用了另一套 reasoning setting。同一模型换档位分数就能大幅移动,所以任何跨系统比分表,都得先标注 effort 档位、观测格式和评测集,再谈差值。
In the AVO configuration, the LLM operated in a text-only modality: each observation was supplied as an exact 64 x 64 text grid, with no images or image tokens sent to the model.
与 VISTA 对照时的关键混杂项:对方主配置喂 512×512 渲染图,AVO 直接喂 64×64 文本网格。感知通道都换了,动作效率自然不可比。要警惕把“观测表示的胜利”读成“智能体架构的胜利”——真正的运行单元是模型×harness×工具环境×上下文策略。
Kahn, R., Kennedy-Shaffer, L., Grad, Y. H., Robins, J. M., & Lipsitch, M. (n.d.). Potential Biases Arising from Epidemic Dynamics in Observational Seroprotection Studies. American Journal of Epidemiology. https://doi.org/10.1093/aje/kwaa188
Carozzi, F., Provenzano, S., Roth, S. (2020). Urban Density and Covid-19. Retrieved from http://cep.lse.ac.uk/pubs/download/dp1711.pdf
Adelani, D. I., Kobayashi, R., Weber, I., & Grabowicz, P. A. (2020). Estimating community feedback effect on topic choice in social media with predictive modeling. EPJ Data Science, 9(1), 1–23. https://doi.org/10.1140/epjds/s13688-020-00243-w
Monforte, A. d’Arminio, Tavelli, A., Bai, F., Marchetti, G., & Cozzi-Lepri, A. (2020). Effectiveness of hydroxychloroquine in COVID-19 disease: A done and dusted deal? International Journal of Infectious Diseases, 99, 75–76. https://doi.org/10.1016/j.ijid.2020.07.056
Pandemic Meets Pollution: Poor Air Quality Increases Deaths by COVID-19. COVID-19 and the Labor Market. (n.d.). IZA – Institute of Labor Economics. Retrieved July 31, 2020, from https://covid-19.iza.org/publications/dp13418/
Urban Density and COVID-19. COVID-19 and the Labor Market. (n.d.). IZA – Institute of Labor Economics. Retrieved July 30, 2020, from https://covid-19.iza.org/publications/dp13440/
Maarten van Smeden on Twitter: “This is a kind reminder that most issues with data (e.g. measurement error, incomplete data, confounding, selection) do not disappear just because you have N = ginormous” / Twitter. (n.d.). Twitter. Retrieved July 19, 2020, from https://twitter.com/MaartenvSmeden/status/1283313496382373890